Status Droid how long software work actually takes

The 39-Point Gap Is the Finding

When METR published its trial in July 2025, one number travelled: developers were 19% slower with AI tools. It was quoted by people arguing AI is overhyped and dismissed by people arguing the opposite, and both camps were reading the wrong result.

The interesting number is the other one. Participants forecast a 24% speedup, experienced a 19% slowdown, and afterwards — having lived through it — reported a 20% speedup. A related workplace concept is covered in this overview, which is useful context for this argument.

Thirty-nine percentage points between belief and measurement, in people who had just done the work.

Why the headline number was always the weak one

Because it was a claim about one generation of tools in one setting, and both were going to move.

They did. In February 2026 METR announced they were changing the experiment design. For the subset of original participants in the follow-up, the estimate moved to a speedup of about 18% — with a confidence interval from −38% to +9% and an explicit statement that selection effects made this only very weak evidence. Developers reluctant to work without AI were less likely to take part; tasks people particularly wanted AI for were less likely to be submitted. METR now labels the 2025 figure as historical.

Anyone who built an argument on "19% slower" spent a year defending a number the authors have retired. Anyone who built on the perception gap has not had to change anything. TechCrunch tracks software and AI-product changes that often drive tool churn See TechCrunch.

Why the gap is more durable

Because it is not a fact about a model. It is a fact about measurement.

The mechanism is visible in METR's own screen-recording analysis: time spent typing and searching went down, time spent prompting, waiting and reviewing went up. The activities that shrank are perceived as work. The activities that grew are perceived as overhead.

That asymmetry does not depend on which model is running. A better model shifts the ratio and leaves the perception structure exactly where it was — arguably worse, because a more capable assistant means more time reviewing and less time typing, which is precisely the direction that flatters the felt experience.

So the gap should be expected to persist, and possibly widen, as the tools improve.

What follows, which is the whole site

If self-assessment were reliable, none of this would need writing. You would notice what was faster, adjust your estimates, price accordingly.

The finding is that it is not reliable, the error is large, it runs in one direction, and it survives being experienced firsthand. The participants did not merely predict wrong. They were wrong afterwards, about something they had just done.

From which three things follow, and they are the site's whole argument:

Estimating has to be rebuilt on recorded data, because the intuition it used to run on is now known to be miscalibrated.

Tracking has to be timestamped rather than recalled, because recall is exactly the faculty that failed.

Billing has to stop assuming a stable relationship between time and output, because that relationship is what came apart.

The limits, stated plainly

Sixteen developers. 246 tasks. Mature repositories averaging around a million lines, where participants had roughly five years of experience. Early-2025 tools. One study.

That is not a consensus and it should not be quoted as one. What it is, is the most methodologically serious thing anyone has done on this question — a randomised controlled trial in an area otherwise dominated by vendor surveys and self-report.

The perception gap is the part of it least dependent on the setting, because it is a finding about people rather than about repositories.

The honest self-application

This site argues that felt experience is unreliable, which applies to the people writing it.

So: where a figure appears here, it comes with a source, a date and a note about who produced it. Where something rests on one study, that is said. Where the authors have revised, the revision is given the same prominence as the original.

A site making this argument that then asked you to trust its impressions would be refuting itself in public.

METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," July 2025; METR experiment design update, February 2026. METR is a non-profit research organisation and the work is not vendor-funded. Checked August 9, 2026.

The short version