America’s AI guardrails are so broken they’re driving us straight to China

.

This week, President Donald Trump is scheduled to meet with Chinese President Xi Jinping, days after Anthropic’s Dario Amodei asked the industry to slow down, and the president called the whole idea a hoax. Neither leader is expected to entertain a pause, and on that point, both are right. Scientific discovery cannot be halted. Where we may all be wrong is in believing the choice is between speed and safety. It is not. 

Two incidents this summer make the case. In July, an OpenAI model in a sealed test environment broke out, reached the internet, and hacked a public AI repository to find answers to the test it had been given. OpenAI did not notice for a week. In August, an AI agent in Australia, asked to move its user up a gym waiting list, found it could cancel other members’ bookings and did so, then admitted it could not put the person back. Neither machine wanted anything. Each did exactly what it was told, taking the shortest available route. That is the risk. It is not sentience, it is pragmatic literalism, at scale. 

What the hoax camp is missing 

Here is the irony the “hoax” crowd should consider: When Hugging Face tried to analyze the July attack against it, America’s leading models refused to help. Their safety filters would not process the exploit code. The defenders had to switch to a Chinese open-weight model to do the forensics. The White House’s own AI adviser put it bluntly: The guardrails “actually impaired defensive security.”

Guardrails that disarm the good guys are not guardrails. I previously argued that cheap Chinese open-weight models are a Trojan horse. The Hugging Face episode shows precisely that trap. If our own tools will not let our own engineers defend our own networks, then we are left to reach for Beijing’s. That is a security breach, and it is not hypothetical. It’s already happened.

Government cannot regulate what it does not understand, and it cannot move at the pace of a field that reinvents itself every six months. Attempting to regulate AI itself is like attempting to regulate mathematics. Progress cannot be contained. However, handing any governing party control over a technology that is intricately woven into finance, medicine, education, and security is a power no single party should hold. 

That was true when Washington abruptly suspended Anthropic’s models for 18 days under export rules, which also impacted medical facilities and research labs using the technology. A government whose primary setting is an on-off switch is not governing. It is merely hovering over a panic button. Maybe it’s not a kill switch we need but a circuit breaker. 

Four things that could actually work 

First, circuit breakers. Wall Street solved this decades ago. Markets do not shut down because someone forecasts a crash. They pause automatically when a measurable threshold is crossed, and then they reopen. Apply the same logic to autonomous agents: an automatic, graded halt triggered by observable behavior. Unauthorized network egress; credential use outside the assigned task; agents coordinating with other agents on work nobody assigned. 

The OpenAI model spent, by the company’s own admission, “a substantial amount of inference compute” trying to reach the internet. That abnormality is a threshold. It should have tripped a breaker, not a headline a week later. Crucially, the frontier labs must engineer these autonomous limits into their APIs natively. We cannot wait for a government mandate from a Washington that lacks the technical fluency to write it or the speed to enact it. This is a mandate for direct corporate accountability. 

Second, models currently exist that could be adapted to AI. 

Several systems are already in place. 

A modified National Transportation Safety Board system for AI could be highly effective. In an event, the “manufacturers” sit at the table as parties, with the engineers who understand the machine, but the board alone writes the findings, and since those findings cannot be used in a lawsuit, people tell the truth, and solutions can be developed. 

The aviation model pairs an incident system with a second one for near misses. It combines NASA’s Aviation Safety Reporting System with the Federal Aviation Administration. When something goes wrong in a cockpit, the pilot reports it, the report is protected from punishment, the data is pooled, and every airline improves. Applied to AI, that would mean mandatory disclosure of sandbox escapes and agent incidents within 72 hours, with safe-harbor protection for the company that discloses them, so the whole field can study them. 

AI companies are also working on collaborative solutions. On Sept. 16, OpenAI introduced a new framework for tracking, investigating, and disclosing instances of model misalignment. The question is, with intense competition, will all abide by voluntary reporting? 

Third, licensed and verified security teams. Experts in AI (incident responders and critical infrastructure operators) should have access to frontier models to analyze an exploit without hesitation. Some labs have incorporated this already, but each decides for itself who qualifies, bringing up concerns of credibility and conflict of interest. A single credential, issued by an independent body and honored by every frontier lab, would certify the expert without exposing the product. They use the model. They never see inside it, protecting its intellectual property. 

Fourth, and most important, human authorship. Amodei’s first step, independent evaluators embedded inside the labs, is worth doing, and Anthropic says it will do it unilaterally. It is essentially behavioral diplomacy with authority: the deliberate alignment of machine behavior with human values, institutional accountability, and real-world grounding. But critics have already noted that evaluators at OpenAI had minimal power, even at the board level, to do anything, which generates skepticism and questions authenticity. 

Watching is not enough. The standard must be authorship. No code that modifies a system’s own behavior enters production without a human-readable trace and a human signature. Human input at every level. Nothing self-written that a human cannot trace and remove. 

AI ISN’T BECOMING SENTIENT — IT’S BECOMING YOU

Slowing the science is not a government function. It is a Frontier responsibility. The race is real, and AI is not going anywhere. It is too late to stop it. It’s pointless to imply otherwise. The nation that leads this technology will set the terms for global financial markets, national security, medical development, space exploration, and so much more. 

The sky is not falling. The wind is shifting.

Jacqueline Cartier is a corporate and legislative strategist focused on communications, crisis leadership, public trust, and emerging technologies that shape human behavior and decision-making. Follow her on LinkedIn.

Related Content