The Agents Hit 66%. The Coordination Layer Is Still You.
Last year, AI agents succeeded at 12% of real computer tasks.
This year: 66.3%. Within six percentage points of human performance. Stanford’s 2026 AI Index published those numbers this week, and the trajectory they imply is not subtle. On software engineering benchmarks, agents moved from 60% to near-100% in a single year. On real-world task completion across domains, 20% to 77.3%.
The agents got capable. Faster than most people expected.
But the question nobody’s asking: who’s managing them?
The benchmark that got measured
Stanford can measure agent task success because the task is defined, the success criteria are clear, and you can score it. 66.3% on OSWorld. 77.3% on Terminal-Bench. These are real numbers from real evaluations.
The harder question is what you can’t put a number on: the system above the agents.
97% of organizations are now actively exploring agentic AI. Only 12% have a centralized platform to understand and control what their agents are doing. 36% have no formal supervision plan at all. 67% of executives believe their company has already suffered a data leak from an unapproved AI tool running somewhere they didn’t know about.
The capability gap closed in one year. The coordination gap didn’t move.
That asymmetry is the actual story.
What changes at 66%
There’s a model most organizations are running — implicitly, because nobody designed it — for how humans relate to AI agents. Call it the supervision model: human reviews agent output, catches errors, approves results, re-runs if wrong.
This worked when agents were at 12%. At 12%, agents fail often and obviously. The human is faster than the agent on anything non-trivial. Supervision is the right frame.
At 66%, the model starts to break. The agent isn’t the bottleneck. The person monitoring it is.
BCG’s research on “AI brain fry” quantified part of this: workers in high-oversight workflows — reviewing, correcting, interpreting model outputs — reported 14% more mental effort and 19% greater information overload compared to workers using AI for direct task replacement. Oversight isn’t free. At 12% agent capability, the overhead was worth it. At 66%, you’re spending significant cognitive energy supervising a system that’s right most of the time — and the cognitive cost compounds with every additional agent.
The deeper problem: at 66%, the agents’ outputs need to be reconciled with each other. Your writing agent produced a draft. Your research agent found contradicting data. Your meeting agent captured a commitment that conflicts with what the writing agent assumed. All of that reconciliation work falls to one person: you.
Not because you’re the supervisor. Because you’re the only thing connecting them.
The human is the coordination layer
This is the pattern across every piece of enterprise AI adoption data right now.
BNY Mellon deployed 20,000 AI agents across its global workforce — one of the most ambitious enterprise AI rollouts to date. The architecture they built: a multi-agent orchestration layer (Eliza 2.0) where agents have their own logins, their own roles, and human managers. They understood that deploying agents without the layer above them creates what every other organization now has: sprawl without accountability.
Most organizations didn’t build that layer. 97% deployed agents. 12% built the infrastructure above them. The remaining 85% have human workers serving as the de facto coordination layer — re-establishing context between systems, hand-carrying commitments from one tool to another, remembering what was decided where.
This isn’t a transitional state. For most knowledge workers, it’s the permanent state. And it gets harder as agents get better, not easier.
The governance policy problem
Every industry report on agentic AI converges on the same recommendation: better governance. Clearer policies. Oversight frameworks. Centralized IT control.
That’s the wrong prescription. Not because governance doesn’t matter — it does — but because policies don’t close loops. A governance framework tells you which agents are authorized. It doesn’t track what they committed to. It doesn’t surface what they left unresolved. It doesn’t reconcile the output from your research agent against the output from your writing agent against the commitment you made in last Tuesday’s meeting.
Policy is a rule about what agents can do. Coordination is infrastructure that knows what they did, what’s in progress, and what still needs to happen.
At 12% agent capability, you could get away with policies. At 66%—and rising—you need infrastructure.
The benchmark that doesn’t exist
Stanford published the agent capability score. That’s the benchmark that’s easy to measure: controlled task, defined success criteria, clean score.
Nobody published the coordination layer score. Not because it’s impossible to measure, but because the coordination layer doesn’t exist in most organizations to measure.
The 2026 AI Index tells you the agents are approaching human performance on real computer tasks. What it doesn’t say — because there’s no data — is whether the system above the agents is keeping up.
It isn’t. The numbers make that clear. 97% deployment, 12% infrastructure. The gap between agent capability and organizational readiness isn’t closing. It’s widening.
At some point this year, most knowledge work domains will cross the 70% capability threshold — the point where the agent is faster and more accurate than the human reviewing it. At that point, the supervision model doesn’t degrade gracefully. It breaks. The human can no longer serve as the coordination layer without tools specifically built to help them do it.
That infrastructure doesn’t get built by writing better governance policies. It gets built the same way BNY Mellon built Eliza 2.0 — by treating coordination as a layer, not an afterthought.
The agents are ready. That question is answered.
The question worth asking now: what’s managing them?