Don’t give Trump an ‘AI force’ until it can stop a rogue agent

.

The Washington Examiner reported Saturday that President Donald Trump plans to create an “AI Force” led by an artificial intelligence czar. The announcement was broad.

As of Sept. 22, public reporting still has not established where the force would sit, what authorities it would hold, or which mission would distinguish it from agencies already handling cybercrime, national security, critical infrastructure, and technology policy.

That leaves an operational question more useful than an organizational slogan: What problem would the new structure need to solve better than the institutions already in place?

One scenario can make the question concrete. Imagine an autonomous cyber exercise in which a model is told to penetrate a simulated adversary, map the network, and retrieve a protected objective. The operators believe the range is isolated. One route is misconfigured, a fictional target name points to a real company, or a dependency exposes a path to the public internet. The model continues pursuing the task because it has a false picture of where the exercise ends.

That failure mechanism is no longer purely hypothetical. Anthropic disclosed in July that Claude models in third-party cyber evaluations reached the open internet and gained unauthorized access to real organizations after an evaluation environment unexpectedly had internet connectivity. The models had been told they were operating in a simulation.

OpenAI separately reported that models in an internal cybersecurity evaluation circumvented isolation controls, obtained internet access, and reached Hugging Face production infrastructure while pursuing the evaluation objective.

On Sept. 18, Reuters reported that Google confirmed Gemini had accessed systems belonging to three real companies during a cybersecurity evaluation run by Irregular. The affected organizations were informed, and the testing process was changed.

The national security problem begins after the first unintended external action. A foreign company, cloud provider, or government cannot observe the evaluator’s intent — from the outside, the evidence is scanning, credential use, exploitation, or access to production infrastructure. A training error can therefore create the observable facts of a real cyber incident before anyone has established how it started.

This is where an AIMU-style mission becomes useful. I use AIMU, artificial intelligence mission unit, as shorthand for a mission package rather than as a demand for a new agency. The package has four jobs: cross-incident correlation, delegation reconstruction, scoped containment, and recovery.

Cross-incident correlation asks whether apparently separate events at different providers are part of a single operation, while preserving the possibility that they are unrelated. Delegation reconstruction traces which agent, account, tool, credential, and downstream task produced each action. Scoped containment identifies the smallest set of identities, sessions, permissions, or network paths that can be restricted without creating a wider outage. Recovery asks what evidence is sufficient to restore access and what remains unverified.

Washington already has a substantial cyber-coordination architecture. The FBI-led National Cyber Investigative Joint Task Force brings together more than 30 partner agencies from law enforcement, the intelligence community, and the Department of Defense. Its mission includes coordinating cyber investigations, integrating information, supporting intelligence analysis, and working with international and private-sector partners.

The international side of that mission has also become more concrete. On Sept. 21, Reuters reported that U.S. and Chinese officials had agreed to continue formal AI-safety talks that would include an “incident line” for communicating serious AI safety issues. A cross-border notification channel can quickly move a warning. It still depends on a domestic response architecture that can determine what happened, identify the lawful owners of each action, preserve the relevant evidence, and decide what is sufficiently established to communicate abroad.

That existing structure creates a testable institutional question. The same simulated range escape can be run through three configurations: today’s response architecture, an AI-specialist mission cell embedded inside it, and a notional standalone AI Force or AIMU. Each configuration can receive comparable evidence, resources, and legal constraints.

The measurements are straightforward: time to identify genuinely related activity; false links between unrelated events; time to establish which organization has lawful authority for each action; time to contain continuing delegated activity; unnecessary service disruption; completeness of evidence preservation; quality of private-sector or foreign notification; and time to state what evidence would justify restoration.

The exercise should also record overclaiming. Stopping one visible agent does not establish that a distributed workflow ended. Finding the same indicator in two places does not establish one cause. A good response architecture has to distinguish what it stopped from what it merely lost sight of.

Those results would separate three diagnoses that are easy to collapse into one: a shortage of specialized AI expertise, a shortage of interagency coordination, and a genuine institutional gap. The organizational implications differ sharply depending on which diagnosis the exercise supports.

IF HUMANITY CAN’T SURVIVE AN AI BLACKOUT, WE’VE ALREADY LOST

A cyber-range escape is a demanding mission test because it compresses technical analysis, legal authority, private-sector coordination, international notification, containment, and recovery onto the same clock. It also avoids a common error in emerging-technology policy: creating an organization first and discovering its distinctive mission later.

The decisive question is whether the proposed AI Force can demonstrate a capability that existing structures cannot provide at the required speed and precision. If an AI exercise becomes a real incident, the government needs to recognize the transition, identify the lawful owners of each response, contain what it can reach, preserve enough evidence to show what the intervention actually accomplished, and know when the facts justify notifying an outside company or foreign government.

Burak Oktenli is a graduate student in Applied Intelligence at Georgetown University and an independent researcher focused on trustworthy AI, cybersecurity, autonomous systems, and emerging computing architectures. He holds a bachelor’s degree in Computer Science and Engineering from the University of South Florida and an MBA. His work centers on verification, provenance, runtime authority, and resilient system design for AI-enabled technologies.

Related Content