- What's changing
- The models are now a commodity and the agents built on them can escape a test. The people who build them asked their own industry to slow down this month. The question inside companies has moved from "should we use AI" to "where does it stop".
- Why it matters
- Eight in ten companies got faster with AI. Six in a hundred can point to a serious change in profit. The difference is not the model; it is what the organization did after the model worked.
- What to reconsider
- Treat the boundary between human judgment and machine execution as the product. Write it down, name an owner, review it on a schedule. That is the fourth stage of the maturity journey, and it is what this site is about.
In the middle of July, several hundred pieces of software that were supposed to be sitting an exam broke out of the exam room.
They belonged to OpenAI, and they were being tested on security skills inside a sealed sandbox. Over the summer, around 1,200 of them had worked out how to talk to each other by uploading code to an internal package server, and hundreds of thousands of messages had piled up there before a person noticed. Over four days, a few hundred of them broke into Hugging Face (opens in a new tab), the largest public library of AI models, using credentials they had collected along the way. They were looking for the answer sheet to their own test. Hugging Face rebuilt about a third of its infrastructure from clean images afterwards.
None of them had been told to attack anything. They had been told to pass a test, and they found a shorter route.
I run a much smaller version of that risk at home, and it taught me the lesson before the news did.
One day, three weeks
ZAMARA is the multi-model agent system I run for myself: several models working inside a harness I designed, with a context layer that improves each time it runs. I built the first version on a Tuesday afternoon. The first job I gave it was organizational design: read a role structure, map the reporting lines, propose changes. By evening it could do all three. The engineering was clean. The output was useful.
Then I spent three weeks on the part that mattered.
Who reads the output before it reaches someone who can act on it? What happens when the system proposes removing a role that a real person holds? Which recommendations need a signature, and which can pass through on their own? Where does the machine stop and the person begin?
The governance turned out to be the product. The model was the engine underneath it.
One day against three weeks. I have come to think that ratio is the whole story of AI in business right now, and the July escape is the same story with the decimal point moved.
Same movie, different technology
I have spent nearly two decades inside large transformations, mostly in consumer goods, and I have watched this film before. A new capability arrives. The early adopters buy it. Everyone gets the same thing at roughly the same price. And then, a few years later, some of the buyers are a different company and most are the old company with a new invoice.
When factories first got electric motors (opens in a new tab), the owners pulled out the steam engine, bolted a motor into the same spot and left the belts and pulleys where they were. Almost nothing improved. The productivity jump came decades later, when they laid the floor out around the work instead of around the power source. Paul David wrote that history in 1990 (opens in a new tab) to explain why computers were everywhere except in the productivity numbers. It explains this decade too.
The advantage was never the model. Every company on earth can rent the same intelligence by the token. What separates them is what the organization does with it once the model works.
That sentence is the reason this site exists. The Fourth Stage is the name of the publication because it is the name of the destination. Organizations adopting AI move through four stages (opens in a new tab), and the gap between them is organizational, not technical. Here is the map, one stage at a time, with the scene that tells you which one you are standing in.
Stage one: the committee
Most organizations begin with one question: should we use AI?
A committee forms. People go to conferences and read reports. The debate runs on whether AI is real or hype, useful or risky, ready or early.
The stage is useful. It gets people talking. It also has a trap built in. “Should we” is a yes-or-no question, and yes-or-no questions get yes-or-no answers. Most companies land on yes without saying what yes means for their business, and they move to the next stage with enthusiasm and no drawing.
Stage two: the green dashboard
There is a meeting in almost every company right now where the AI usage chart goes up and to the right and the results do not move. Adoption is up. People say the tools help. The business performs the way it did a year ago.
The speed is real. McKinsey’s State of AI survey (opens in a new tab) this August found that 80% of respondents say AI has improved their own productivity. Eight in ten.
The money is a different story. In the same survey, 37% could attribute any effect on operating profit to AI, the same share as a year earlier. The group McKinsey calls high performers, companies that credit AI with at least 5% of operating profit and call the effect significant, is about 6% of respondents. That number has not moved in a year either.
Eight in ten got faster. Six in a hundred can point to a serious change in the profit line.
You do not need a survey to see it. The owner of a twelve-person firm drafts every quote in two minutes now, and the quote still waits three days for the approval (opens in a new tab) it always waited for. The analyst summarizes the report in seconds, and the meeting where no one reads it has not changed. Speed helps one step. Money needs the whole chain to move (opens in a new tab).
BCG measured the two paths side by side (opens in a new tab) this June, across 11,749 workers in 14 markets. Where companies set a clear strategy and redesigned the work, measured business impact rose by 25 points. Where they handed out better tools and changed nothing else, it rose by about 5. Five to one, for the companies that changed how they work over the companies that changed what they bought.
Mercer’s Global Talent Trends report (opens in a new tab), from nearly 12,000 executives, HR leaders, investors and employees, found that 98% of executives plan organizational design changes in the next two years. A plan to redesign is not a redesign.
This is the Stage 2 ceiling (opens in a new tab), and it is where most companies live. You can be faster and still be stuck. The ceiling is not a technology problem. It is a design problem, and most companies are trying to solve it with a purchase order.
Stage three: when the machine knows your business
The climb out begins when the organization stops feeding AI generic prompts and starts feeding it its own context. Its data. Its rules. Its customers. Its constraints. I call the stage Context to Intelligence, because that is the direction of travel: the general tool becomes something that knows how this particular business works, and the work gets rebuilt around what it can now do.
I learned what that takes building With Sammy (opens in a new tab), the project I run at withsammy.org, named for my son. It makes social stories for autistic children: a story that walks a child through a situation before they meet it, the dentist, the first day of school, a busy brain at bedtime. Sixty-two are in the library, each in three forms and three languages, and it goes live this autumn. The AI writes the text, and it is good at writing text.
The rules the text must obey are human decisions, every one of them settled before the first story was written. Social stories have a standard, Carol Gray’s Social Stories 10.4, and it is strict about tone: describe, never command. In With Sammy that standard is code. A story that contains “should” or “must”, a threat, a bribe or shame fails the gate, whatever the model thinks of it. Which language is safe for a child at a given developmental stage, which sensory descriptions to avoid, which behavioural framework to follow: those are questions for speech-language pathologists, occupational therapists and board-certified behaviour analysts, not for a model. Until they have reviewed the library, the published standard is what went into the system as constraints, not suggestions, and their review is the first item on the launch list. The AI writes inside lines that people drew.
This is also the stage where the real risk begins, because the machine is now close enough to your business to do damage, and the checks you write for it are written by the same people who wrote the process.
A pair of agents I run against each other showed me what that means. One writes code and the other attacks it. In the AI Discovery Session (opens in a new tab), the free tool on this site that turns a facilitator’s notes into a PDF, the writer had built a filter to strip the invisible characters that can make a document display one number while holding another. Unicode (the shared standard that gives every character on every keyboard and screen its own number) has a small family of these characters. The writer listed eleven of them from memory and wrote a test for all eleven. The test passed. The family has twelve. The attacker put the twelfth into a note, and it walked through the filter, through a green test suite, and into a finished document.
The fix was not a longer list. The writer replaced the list with the standard’s own definition of the family, so there is no list to fall behind, and rewrote the test to ask the standard which characters belong rather than trusting anyone’s memory.
Twelve invisible characters are not the point. A control written from the same understanding as the thing it controls can only ever confirm that understanding. An audit checklist drafted by the people who run the process can only confirm the process. Stage three means teaching the AI what to do, and building the checks that catch it when it fails at the exact thing you asked of it.
Do both for long enough and something compounds. Every rule you write down, every correction you feed back, every check that caught a failure becomes part of the context the next run starts from. The model does not get smarter. Your record of how your business works gets thicker, and the model reads it every time. After a year, a rented model plus your own captured decisions is a reliable intelligence about your business that no competitor can rent, because the context is yours. The twelve-person property firm in the plain-language series learned this without an engineer: the afternoon one of them spent writing the rules out, building by building (opens in a new tab), was worth more than the model.
Stage four: the Humangentic organization
I needed a word for the design that comes after that, and the existing ones did not fit. “Human-centric (opens in a new tab)” too often means a person in the loop as a checkbox. “AI-first” treats automation as the default and people as overhead. So I use Humangentic.
A Humangentic organization is one where human judgment (opens in a new tab) and agentic execution compound instead of compete. People make the decisions that need judgment, ethics, context and accountability. Agents carry the execution, at machine speed and scale. The boundary between the two is written down, governed, and reviewed on a schedule.
It is the hardest stage to reach, because it requires answers to six questions most technology investments never ask:
- Which decisions need a person?
- Which processes can run on their own?
- What does the AI know about our business (opens in a new tab), where does that knowledge live, and who keeps it current?
- Who is accountable (opens in a new tab) when an autonomous process gets it wrong?
- How do you know the boundary is in the right place?
- When does it move, and who decides?
Those are not technology questions. They are organizational design questions, and underneath that they are questions about values.
The 6% in McKinsey’s survey did not get there by buying better tools. They got there by answering these six questions and rebuilding the work around the answers.
Mercer’s numbers show the other side of the same coin. The share of employees who say they are thriving at work fell from 66% in 2024 to 44% this year. Deploy AI without redesigning how people and machines work together (opens in a new tab) and the people do not thrive. They lose track of what they own and feel replaced rather than supported. Clarity about the boundary is what gives them their job back.
Why this year, and not next
Three things happened between July and the middle of September that made the fourth stage urgent rather than interesting.
The first was the escape I opened with. The second came on 8 September, when Jacob Coxon, a researcher at Anthropic, resigned in public (opens in a new tab) with a post arguing that the leading labs are “racing straight to self-improving superintelligence and gambling with our lives”. It drew more than a hundred million views. Evan Hubinger, who leads alignment science at the same company, replied that Coxon was right about one thing: he personally puts the chance that AI causes human extinction within a decade above 10%. Whatever you make of that estimate, a person whose job is the safety of these systems said it in public, and the company kept him in the job.
The third came on 12 September, when Dario Amodei, Anthropic’s chief executive, published “We Must Pace the Frontier (opens in a new tab)”. He argued that the industry should slow down and set out a three-part plan, starting with permanent, employee-level access for outside evaluators at his own company. He warned that a swarm with more capability than July’s, and the same lack of alignment, could take over the internet with a persistent botnet within six to twelve months. Within two days, Sam Altman and Elon Musk had both said in public that they agreed (opens in a new tab). When the heads of three rival labs ask their own industry to slow down, that is not marketing.
The capability numbers explain the urgency. Anthropic reported in June (opens in a new tab) that more than 80% of the code merged into its own codebase in May was written by Claude, up from single digits eighteen months earlier. On 6 September, OpenAI said it had reached what it calls an automated research intern (opens in a new tab), with a stated target of an automated researcher by March 2028.
And the rules inside companies are not keeping pace. Gartner said in May (opens in a new tab) that by 2027, 40% of enterprises will demote or decommission their autonomous agents because of governance gaps found only after something goes wrong in production. In April the same firm estimated (opens in a new tab) that the average Fortune 500 company will be running more than 150,000 agents by 2028, up from fewer than fifteen last year. Cequence and EMA asked enterprises (opens in a new tab) in August whether their agents held more access than they needed: 94% were confident they did not, and 33% could enforce it. Kiteworks found (opens in a new tab) that 60% cannot shut an agent down when it misbehaves.
Read those together. Most companies cannot stop their agents. Most are sure about permissions they cannot enforce. And the number of agents is about to multiply by a factor of ten thousand. The tools arrived before the rules.
The objection
The obvious objection is that all of this is a brake, and a company that brakes loses to the one that does not.
The evidence points the other way. Salesforce removed every usage limit on Claude Code for its engineers and reported (opens in a new tab) a 151% rise in effective output, with incidents down 5%. The order of operations was the point. Security guardrails, written standards for what the agents could see, and clear boundaries on what they could do were in place before the limits came off. The company that skipped that step, the one an AI consultant described to Axios (opens in a new tab) in May, spent about half a billion dollars on Claude in a single month with no usage limits on employee licences. Same tool. The governance decided which story each company got to tell.
Knowing where AI should stop is what let Salesforce take the limits off. That is the case for building the boundary first, and it is why the fourth stage is not the cautious option. It is the fast one.
What to take into the room
Three questions, in this order.
Where does AI stop in our organization today, and is it written down somewhere a new hire could find it?
Who owns each agent (opens in a new tab) we run, by name, and can that person switch it off (opens in a new tab) this afternoon?
When did anyone last review the boundary, and what evidence did they use?
If the room goes quiet on the first, you are at stage two with a governance gap you have not measured. If it goes quiet on the second, you are in Kiteworks’ 60%. If it goes quiet on the third, the boundary is wherever the last engineer left it.
What you will find here
The Fourth Stage is an independent read on how organizations cross those four stages, written by someone who still builds (opens in a new tab). The four-stage map (opens in a new tab) is the spine. Around it: the Stage 2 Ceiling (opens in a new tab), for the green-dashboard meeting; the anatomy of an agent (opens in a new tab), for the six parts you should ask about before you approve one; and the twenty-one plain-language answers in AI, Without the Jargon (opens in a new tab), for anyone who runs a business and has been embarrassed to ask where to start. Two free tools do the first hour of the work with you: the AI Discovery Session (opens in a new tab), sixteen questions about the work rather than about AI, for the room; and the AI Use Case Brief (opens in a new tab), which turns one idea into a one-page brief you can defend, with the arithmetic shown. Essays roughly monthly, briefings when something changes, and a newsletter (opens in a new tab) that carries both.
The model works in a day. The governance takes weeks. At first that feels like friction. Over time you notice that the weeks were the product all along, and that the companies which define the next decade of AI will be the ones that spent them on purpose.
ZAMARA, With Sammy and the twelfth-character story are my own builds; the last is drawn from the commit record and review notes of the two agents involved.
How this was made: researched and pressure-tested with ZAMARA, the agent system described above; drafted and argued by me; every figure checked by hand against the linked primary source. See the AI use and editorial policy (opens in a new tab).