METR’s randomized controlled trial is one of the more uncomfortable data points in the AI productivity conversation, and one of the most useful. Sixteen experienced open-source developers, averaging five years on the mature repositories they worked in, completed 246 real tasks, each randomly assigned to allow or disallow AI tools. Before starting, the developers forecast that AI would cut their completion time by 24 percent. After finishing, they estimated it had cut their time by 20 percent. The measured result: AI tools made them 19 percent slower.
That gap, between what experienced engineers believed and what actually happened, isn’t an argument against AI-assisted development. It’s the sharpest available evidence that writing code with AI assistance and supervising AI output effectively are two different skills, and most engineers have only been trained in something adjacent to the first one.
Why Experienced Engineers Got Slower, Not Faster
The developers in METR’s study weren’t AI skeptics forced into an unfamiliar tool. They had tens to hundreds of hours of prior LLM experience and used state-of-the-art tools for the period, primarily Cursor Pro with Claude 3.5 and 3.7 Sonnet, on codebases they knew well. The slowdown showed up specifically on real, complex, context-heavy work, not the kind of greenfield task AI demos are built around.
The gap between forecast and measured outcome is the part worth sitting with. If the friction were simply that AI writes code slowly, developers would have felt the slowdown in real time and adjusted their estimate downward. They didn’t. They finished the study still believing AI had saved them time. That disconnect points to where the actual cost was hiding: not in code generation, but in the reviewing, verifying, and correcting work required to trust AI output on a codebase with real history and context, work that doesn’t register the same way in a developer’s sense of their own productivity.
By early 2026, METR had to change its own experiment design, noting that as agentic tools like Claude Code and Codex became standard in the developer population it was recruiting from, finding a comparable group willing to work without AI access had become difficult. Even the research measuring this shift is being reshaped by how fast the baseline is moving.
The Skill That's Shrinking vs. The Skill That's Growing
Writing code from scratch was already a smaller share of a developer’s week than most organizations assumed before agentic AI entered the picture. What’s changing now is what fills the rest of that time: less original code production, more reading, reviewing, and verifying code someone, or something, else produced.
McKinsey’s research on the agentic organization describes this as humans moving “above the loop,” from executing every step of a task to steering outcomes and evaluating what an agent has already done, stepping back into the loop selectively where judgment or direct contact still matters. The three profiles McKinsey identifies emerging alongside this shift, M-shaped supervisors who orchestrate across domains, T-shaped experts who handle exceptions and safeguard quality, and AI-augmented frontline workers, map directly onto what “supervising an agent” actually requires in an SDLC context: system-level judgment, pattern recognition on the edge cases where agents reliably fail, and the ability to catch output that’s subtly wrong rather than obviously broken.
Reskilling Isn't Just a Junior-Engineer Problem
Writing code from scratch was already a smaller share of a developer’s week than most organizations assumed before agentic AI entered the picture. What’s changing now is what fills the rest of that time: less original code production, more reading, reviewing, and verifying code someone, or something, else produced.
McKinsey’s research on the agentic organization describes this as humans moving “above the loop,” from executing every step of a task to steering outcomes and evaluating what an agent has already done, stepping back into the loop selectively where judgment or direct contact still matters. The three profiles McKinsey identifies emerging alongside this shift, M-shaped supervisors who orchestrate across domains, T-shaped experts who handle exceptions and safeguard quality, and AI-augmented frontline workers, map directly onto what “supervising an agent” actually requires in an SDLC context: system-level judgment, pattern recognition on the edge cases where agents reliably fail, and the ability to catch output that’s subtly wrong rather than obviously broken.
What the Reskilling Curriculum Actually Needs to Cover
A handful of specific capabilities separate engineers who get real leverage from AI-assisted development from the ones the METR study describes.
Verification and review discipline has to become an explicit, trained skill rather than an assumed one, treating AI output the way a careful engineer would treat a junior developer’s pull request rather than accepting the first suggestion. Defining intent and acceptance criteria clearly before an agent starts work matters more than it used to, since McKinsey’s research frames the human role specifically as declaring high-level intent and boundaries, then evaluating what comes back. System-level and architectural judgment, the parts of the job METR’s study shows AI struggles with most, become the highest-value skill precisely because they’re the hardest to automate: ambiguous requirements, cross-team tradeoffs, and debugging issues that depend on deep context AI doesn’t have. Knowing when not to trust output matters as much as knowing when to accept it, particularly on unfamiliar or high-stakes surface area, since METR’s finding shows the slowdown concentrated exactly where deep system knowledge mattered most. And escalation literacy, knowing when a task should route to a governance or compliance check rather than ship on an engineer’s judgment alone, becomes part of the job rather than a separate compliance function bolted on afterward.
Why This Has to Be Deliberate, Not Assumed
None of this happens as a side effect of giving engineers AI tools and expecting them to figure it out. McKinsey’s broader research on the agentic organization points to a rule of thumb worth taking seriously here: roughly five dollars in people investment for every dollar spent on the technology itself. Treating reskilling as a byproduct of tool access rather than a funded program is the same mistake in a different shape.
There’s a real tension worth naming honestly rather than smoothing over: Gartner’s research on smaller engineering teams warns that organizations which cut junior hiring in response to AI will hollow out their own talent pipeline by 2028, since junior engineers are where the judgment this whole model depends on has traditionally been built, through the slow, unglamorous work of writing code and getting it wrong. If junior engineers spend less time writing code from scratch, the pipeline that used to produce reviewers and system-level thinkers needs a deliberate replacement, not an assumption that judgment will develop the same way it always has. The market is already repricing around this: McKinsey Global Institute research found demand for AI fluency grew sevenfold in two years, faster than any other skill category, which means organizations that don’t invest in this reskilling deliberately will simply lose the people who’ve built it elsewhere.
Where Ascendion Fits
This is the premise behind Ascendion’s Carbon + Silicon model: the skill of supervising and verifying agent output is treated as the product of the operating model, not a side effect engineers are expected to develop on their own. AAVA™ keeps humans in the loop and accountable by design, and Ascendion’s talent orchestration capability, METal, is built to match engineers into the supervision, verification, and orchestration roles this shift actually requires. The gap in METR’s study, between engineers who believed AI had made them faster and the measured reality that it hadn’t, is exactly the gap a deliberate reskilling program is meant to close.
Supervising agents effectively is a trained skill, not a byproduct of tool access. See how AAVA builds human accountability into every agentic workflow.
Ascendion is the AI-native disruptor reinventing how global enterprises build software for impact. Its engineering teams, powered by AAVA, the company’s proprietary agentic AI platform, deliver measurable business outcomes: accelerating growth, unlocking capital, and de-risking transformation. With 11,000+ engineering professionals and 12,000+ AI agents working across 12 countries, Ascendion delivers the promise of AI to more than a third of the Fortune 500. Learn more at https://www.ascendion.com.
Engineering to the Power of AI™, AAVA™, and Engineering to Elevate Life™ are trademarks or service marks of Ascendion®. AAVA™ is pending registration. Unauthorized use is strictly prohibited.