Picture two agents inside the same bank. One resets forgotten passwords. The other changes overdraft limits. In most enterprises today, both run under a single rule: locked down together, or trusted together. One policy. Two completely different risks.

That flat rule is not a detail. Gartner’s finding from May names binary governance, agents treated as either locked-down or fully-trusted, as a root cause of enterprise agent failure. That’s the observation. Here is the inference: the binary fails in both directions.

Lock everything down, and your agents can’t do anything worth deploying them for. The program stalls, and the value case quietly dies in committee. Trust everything, and the rule you wrote for the password desk is now also the rule for the overdraft desk. Same policy. Opposite failure.

The alternative I argue for is graded control: the leash sized to the task’s risk. A leash is not a yes-or-no device. It has a length. Nobody walks every dog on the same one, yet that is exactly how most organizations are walking their agents.

How to grade a task

My recommendation: before an agent gets autonomy over any task, grade the task rather than the agent, on three questions.

  1. Blast radius. If this action goes wrong, what does it touch? A botched password reset annoys one customer. A botched overdraft change moves money.
  2. Reversibility. Can the action be undone cheaply, or does the damage land before anyone sees it?
  3. Accountability. Who owns the outcome when the agent acts? If no one can answer, the leash is already too long.

Small blast radius, easily reversed, clearly owned: long leash: delegate it fully and stop reviewing routine wins. Large blast radius, hard to undo: short leash: approval gates, spend limits, or no autonomy at all. Most tasks sit between the poles, and the grading is the point: it forces the delegation decision to be made task by task instead of once, badly, for everything.

This is judgment-layer work

Sizing each leash is precisely what the judgment layer (opens in a new tab) exists to do: deciding where the machine is trusted, where it’s checked, and where it’s forbidden. Delegation is not one decision; it is a portfolio of decisions, one per task. And graded control is what the gates in the anatomy of an agent (opens in a new tab) look like when they are designed rather than defaulted.

What to take into the room: ask how many distinct levels of agent autonomy your governance actually recognizes. If the honest answer is two, allowed and blocked, you are running the bank where the password desk and the overdraft desk share one rule. Grade the tasks. Size the leashes. Then scale.