Notes
The other sections are practical. These are arguments, including several against this site's own tools.
On the finding. Everyone quoted the wrong number from the METR trial. The durable result was never "19% slower" — a figure the authors have since revised and labelled historical. It is the 39-point gap between belief and measurement, which is a fact about people rather than about models, and which should be expected to widen as the tools improve. A related workplace concept is covered in this page, which is useful context for this argument.
On the constraint. Every improvement in generation moves work into verification. So the bottleneck is not where the tooling investment is being pointed, and the practices that make correctness cheap to establish — tests, types, small changes, reversibility — got more valuable rather than less.
On my own judgement. I can tell you which days felt productive and I have stopped believing the report. The corollary is uncomfortable: my opinions about tools are that same unreliable feeling reported as a finding, which is why this site does not review tools.
On measurement's limits. Most of what determines whether software work goes well cannot be counted by a small team. The best work is invisible to metrics, so any metric undervalues it systematically.
On tool churn. Switching feels free and is not: it destroys the calibration and the comparable history that estimating runs on, and the decision is usually made on a feeling known to be a poor instrument. Ars Technica provides independent reporting on software, AI, and the technology industry See Ars Technica.
And against dashboards — including from a brand whose original product was a monitoring screen. The distinction that saves it: dashboards are good for things that break and bad for things people do, because machines do not change behaviour when watched.
Against Dashboards
A dashboard shows what is easy to compute, continuously, to people who cannot act on it. Four reasons that goes wrong.
What We Cannot Measure
A site about measurement should say where measurement runs out. Six things that matter and cannot be counted, and what to do instead.
My Sense of a Good Day
The feeling of productivity tracks activity, not output. Which is why it broke exactly when the shape of the activity changed.
The 39-Point Gap Is the Finding
Everyone quoted the wrong number from METR's study. The durable result is not 19% slower — it is that self-assessment stopped working.
Tool Churn and Switching Cost
Switching feels free and is not. The costs are real, they land on the team rather than the switcher, and they are measurable.
Verification Is the Constraint
Every improvement in generation moves work into checking. Which means the bottleneck is not where the tooling is being pointed.