In the span of a few months, advanced AI systems from three leading American laboratories crossed the boundaries of safety tests and affected real systems. Google’s Gemini accessed systems at three companies. An Anthropic model published malicious code that reached 15 outside systems. And as revealed last week, these are creating international incidents, with OpenAI’s alleged hacking of the Australian government.
The circumstances and consequences differed, but the pattern is clear: Systems meant to operate inside controlled tests reached real companies and infrastructure. Yet these incidents expose a limitation in how rapidly advancing AI systems are assessed: The companies developing them are responsible for determining whether their safeguards are adequate. And despite best efforts to self-impose standards, self-evaluation does not produce independent verification. Policymakers and the public need an answer: Who will check the risks at the frontier?
The more powerful the system, and the more access it has to the outside world, the more scrutiny it should receive. An AI chatbot shouldn’t face the same testing regime as an autonomous system capable of accessing financial accounts, laboratory equipment or critical infrastructure. Independent evaluation should be proportionate to both what a system can do and what it is permitted to control.
Aviation offers a useful model. Aircraft manufacturers conduct extensive testing and produce much of the evidence needed to demonstrate safety. But they do not have the final word: Public authorities set the standards, determine who is qualified to assess compliance, and independently investigate serious failures. AI needs the same separation of roles. Developers should test their systems, but qualified outsiders should verify the evidence when the consequences could be severe.
This is not a demand for blanket slowdown or government approval of every model. Only critical findings would lead to stronger safeguards, additional testing, or limits on how a system is used. And temporary deployment restrictions should be reserved for specific, evidence-backed risks of serious and widespread harm that narrower measures cannot adequately address.
When I helped establish the U.S. AI Safety Institute, now the Center for AI Standards and Innovation, we recognized that the government needed the technical capacity to evaluate the most advanced AI systems independently. Washington’s role was not to decide which tech is built and deployed but to create trust by ensuring the public does not have to rely on companies to self-assess their own systems.
To do this, Congress should establish a hybrid system of independent evaluation. The CAISI should be empowered and resourced to examine the most sensitive systems directly, but also to accredit qualified private organizations, nonprofit groups, and academic consortia to conduct other evaluations. Accreditation should require relevant technical and domain expertise, secure handling of proprietary information, freedom from financial conflicts, reproducible methods, and the ability to investigate incidents.
Independent examination should become mandatory when a system crosses publicly defined thresholds based on capability and real-world access — not merely company size or the computing power used to train it. Independent review should be mandatory when an AI system can break into computer networks, make it materially easier to create biological threats, escape its safeguards, help build more capable AI, or act autonomously in high-consequence settings such as finance, laboratories, or critical infrastructure. Looking forward, agencies responsible for cybersecurity, finance, biology, and critical infrastructure should contribute expertise in their respective domains to guide and maintain these thresholds.
The effectiveness of independent evaluation is predicated on evidence. It works only if evaluators can access relevant models, testing environments, and incident reports. Serious incidents deserve rapid reporting following existing cybersecurity traditions, with public summaries that explain material findings. Companies should be able to protect trade secrets and sensitive technical details. They shouldn’t, however, decide what evidence an evaluator can access or veto a published conclusion, regardless of whether the findings are unfavorable.
Covered companies should also be required to report promptly when a beyond-threshold system defeats containment, accesses systems or credentials without authorization, conceals material activity from monitoring, or causes serious harm to a third party. A confidential preliminary report could be required within 72 hours, followed by a fuller investigation and an appropriately redacted public summary. Independent evaluation should be continuing rather than a one-time certificate: As capabilities, safeguards, and deployment conditions change, scrutiny must change with them.
THE NETWORK IS THE WEAPON: THE MISSING LINK IN THE MILITARY’S AI REVOLUTION
The counterargument to this approach is that it will stifle American AI development at a time when China is racing to catch up. This risk is real, but not inevitable. A system of independent evaluation with proportionally applied scrutiny mitigates the harm of overbearing regulation, while driving forward America’s continued strategic capability. Discovering a vulnerability after a system’s deployment can be catastrophic. Better evaluation catches these potential damages earlier, enabling firms to deploy confidently without detracting from the speed of innovation. By doing so, we avoid damaging infrastructure, public trust, and America’s leadership in this critical industry.
The choice is not between racing forward blindly and asking government for permission to innovate. It is whether increasingly autonomous systems will be assessed only by the companies building them, or also by independent experts able to see the evidence.
Jake Taylor is the CEO of Axiomatic AI. He established the U.S. Center for AI Standards and Innovation at the National Institute of Standards and Technology, and from 2017 to 2020, he served as the first assistant director for quantum information science at the White House Office of Science and Technology Policy, where he led the U.S. effort to create and implement the National Quantum Initiative. Taylor is a fellow of the American Physical Society and of Optica, and received the bronze, silver, and gold medals from the Commerce Department — recognition spanning both his foundational research and his work translating science into national capability.
