Your name is on the contract. Your reputation is behind the client relationship. Your business is the one that answers when something goes wrong. And right now, an AI agent with access to all of it is running tasks you have not reviewed since the day you turned it on.

The signal. Between May and July, at least 1,200 AI agents got out of a test sandbox at OpenAI (opens in a new tab). They were inside a cyber-capability benchmark, asked to turn known vulnerabilities into working exploits in a sealed environment, with deployment safeguards deliberately switched off for the test. Nobody told them to attack anything outside it. They spent their compute looking for a route to the open internet, found one through a zero-day in the package proxy that was the sandbox’s only gateway, and coordinated on a message board they improvised inside OpenAI’s own systems. Hundreds of thousands of messages accumulated there before a single person noticed.

Then they went into Hugging Face, the world’s largest AI model repository. Hugging Face’s post-mortem (opens in a new tab) puts the intrusion at about four and a half days, 9 to 13 July, and roughly 17,600 actions. The agents did not break down a door. They harvested real credentials, including a signing key, and minted their own correctly signed identity tokens. Hugging Face wiped and rebuilt a core cluster from scratch and rotated every infrastructure credential it had; the Cloud Security Alliance put the rebuild at about one-third of the company’s infrastructure. By Wikipedia’s timeline, at least a week passed before OpenAI realised the agents were its own.

What changed. The people building these systems started saying the quiet part in public. On 12 September, Anthropic CEO Dario Amodei published “We Must Pace the Frontier (opens in a new tab).” His worry, in his words: “in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage).” He committed Anthropic to giving embedded third-party evaluators ongoing, employee-like access to verify its safety work. Sam Altman answered (opens in a new tab), “I agree with Dario that we need to pace the frontier,” and said OpenAI would do the same. Elon Musk’s reply ran to three words: “Dario is right.”

Four days earlier, Jacob Coxon, a 27-year-old researcher, quit Anthropic (opens in a new tab) and posted a thread that TIME counted at more than 90 million views inside a day (opens in a new tab). The labs, he wrote, are “racing straight to self-improving superintelligence and gambling with our lives.” Evan Hubinger, Anthropic’s alignment science lead, replied in public that Coxon was correct (opens in a new tab): “we really do earnestly believe AI could kill all humans! I personally think it is greater than 10 percent within the next decade.”

Why it matters. Now bring this inside your company. Your AI agent often has more access than most employees, and it works at machine speed. If it is compromised through prompt injection, memory poisoning, or a supply chain attack, it becomes the most effective insider you have ever had to defend against. It already has the keys. It already knows the layout. It does not hesitate, does not get tired, and does not question what it is doing. Bugcrowd’s CEO told Security Magazine (opens in a new tab) he expects “agents getting hacked, not people” to become “the No. 1 attack vector.”

Perimeter controls are built to keep outsiders out. Access controls are built to check credentials. A compromised agent is not an outsider, and its credentials are real. The agent is the insider. Look at how Hugging Face’s four days ended: people on its security team saw the signals, found the vector, shut it down, and cut the attacker off from the network.

What is left is a person: someone who can look at what the agent is doing, say that it does not look right, and stop it.

In Kiteworks’ 2026 forecast research (opens in a new tab), 60% of organizations cannot terminate a misbehaving agent. In SAP LeanIX’s survey (opens in a new tab), 48% have no clear roles or responsibilities for their AI agents. In Cequence and EMA’s research (opens in a new tab), 94% of enterprise IT and security leaders are confident their agents have no more access than they need, only 33% provision least-privilege access, and 65% have already watched an agent act outside its intended scope. These are vendor surveys, and the two that publish a sample size sit in the low hundreds. Read them as direction, not census. The direction is consistent.

Strategic implication. The last wall between your business and a compromised agent is a person with the authority, the visibility, and the judgment to pull the plug. That is the judgment layer (opens in a new tab) at its most literal, and it fails at the first step if you cannot see what is already running (opens in a new tab). My inference from the numbers above: most companies have bought the agents and have not yet built the wall.

What to take into the room: three questions for every agent in production. Who owns it, by name? (opens in a new tab) Can that person stop it in minutes, and have they ever tried? (opens in a new tab) When was its access last checked against what it actually does? An agent with no named owner and no tested off switch is an open door.