Anthropic Studied Millions of AI Agent Sessions. The 73% Number Tells You Everything.
Anthropic published research last month that analyzed millions of real agent interactions.
The headline most people took: AI agents are getting more autonomous. Session lengths in Claude Code nearly doubled in three months — from under 25 minutes to over 45 minutes at the 99.9th percentile. Among users with experience, over 40% run on full auto-approve, up from 20% for new users.
All true. But that’s not the number that matters.
The number is 73%.
73% of agent actions still have a human in the loop.
Read that again. Among users who’ve gone well past the experimenting phase — who’ve built automations, who’ve granted agents broad permissions, who run multi-hour uninterrupted sessions — nearly three out of four agent interactions still require a human to review, approve, or intervene.
The standard framing for this is a “trust problem.” As in: people don’t yet trust AI agents enough to let them run independently.
That framing is wrong. And believing it leads you toward the wrong solutions.
What 73% actually means
Trust is an outcome. It’s not the thing to fix directly.
You trust a system when it demonstrates it understands what matters to you, when its judgment aligns with yours, and when the cost of being wrong is manageable. You distrust a system when you can’t predict what it will do, when its actions don’t reflect your priorities, and when mistakes are expensive to unwind.
The Anthropic research found that only 0.8% of agent actions are irreversible. Low stakes, mostly. And yet 73% still require human oversight.
If this were a trust problem, you’d expect to see it correlate with high-stakes actions. But it doesn’t. People stay in the loop on routine, low-stakes, easily reversible tasks. They stay in the loop because the alternative — handing control to a system that doesn’t have a model of what you’re trying to do — feels like flying blind.
That’s not distrust of AI. That’s a rational response to missing infrastructure.
The context gap
Here’s what an AI agent typically knows when you hand it a task:
- The task itself (whatever you wrote in the prompt)
- The tools available to it
- Whatever history exists in that specific session
Here’s what it typically doesn’t know:
- What else you’re working on right now
- Which decisions are still open
- What you agreed to with someone last week that might affect this task
- Which commitments you’re actively tracking
- What “done” actually means given the full context of your work
Without that, the agent can execute. It cannot coordinate.
And coordination is the thing that requires judgment. It’s the thing that makes the difference between an action that closes a loop and an action that opens three new ones.
So you stay in the loop. Not because you’re anxious. Because you’re the only one who holds the context. You are the coordination layer — not by choice, but by default.
The efficiency trap
The Anthropic data shows something else worth naming.
As users gain experience, they grant more autonomy. The 40% auto-approve rate among power users looks like progress. But look at what those sessions are actually doing: 49.7% of tool calls are in software engineering. The remainder is distributed across back-office automation, marketing, data analysis, finance.
These are tasks with well-defined inputs, clear success criteria, and contained blast radius. They’re the tasks where handing off works because the scope is bounded. The agent isn’t operating on your behalf across your work — it’s completing a specific, scoped job.
That’s useful. But it’s not what the promise of AI agents actually is.
The real promise is an agent that can handle the ambiguous, the cross-functional, the context-dependent. The thing you’d delegate to a senior operator who knows your situation deeply. That’s not a software engineering task with clear acceptance criteria. That’s coordination work — and for that, the agent still needs you in the loop, because the agent doesn’t know your situation.
Why this is a design problem, not a maturity problem
The conventional wisdom says we’re early. That as agents improve, trust will grow, and humans will gradually step back. The 73% will become 50%, then 30%, then something negligible.
Maybe. But this framing treats the problem as a capability gap when it’s actually a context gap.
The agents aren’t less capable than we need them to be. They’re less informed. The bottleneck isn’t intelligence — it’s context architecture. Agents operate on prompts and sessions. They don’t operate on a persistent, structured model of your work, your priorities, your commitments, and your open loops.
Until they do, you’re stuck being the middleware.
The design question isn’t “how do we make agents smarter?” It’s “how do we give agents a shared model of what matters?” That’s a different problem. It’s an infrastructure problem.
What changes when context is solved
Imagine an agent that, before taking any action, can check it against: what you’re actively trying to close this week, who you’ve committed to and what you’ve committed to, what decisions are still open, and what signals matter to your priorities right now.
Not because you told it these things in the current session. Because the system maintains a persistent, structured model of your work — updated continuously, available to every agent operating on your behalf.
The 73% doesn’t drop because agents become more capable. It drops because agents become more contextual. Because the coordination layer exists as infrastructure, not just as you, in the loop, holding everything together.
That’s what Anthropic’s data is pointing at. Not a trust crisis. A context architecture gap.
The trust will follow when the context is there.
Deeplica is being built as the coordination layer: a persistent AI that holds your context across sessions, understands your open loops, and acts on your behalf across the tools you already use. If the 73% problem sounds familiar, we’d like to hear from you.
Source: Measuring AI agent autonomy in practice, Anthropic, February 2026.