Measuring
The case for measuring is narrow and it is strong: perception of speed is demonstrably miscalibrated, by around 39 percentage points in the best evidence available, in the direction that flatters the tool.
Measurement corrects that one documented failure. It does not tell you whether the work is good, and nothing in this section pretends otherwise. For a practical implementation reference alongside the measurement argument here, see learn more.
What to record is smaller than you think. Five fields per task — timestamps, category, whether an assistant was used, the estimate written down before starting, and a flag for rework. Fifteen seconds. Anything costlier gets abandoned inside a fortnight, and inconsistent data is worse than none because you will trust it anyway.
Granularity is the trap. METR labelled 143 hours of screen recordings to get second-level resolution. That is a research budget. Every decision a small team actually makes — how to estimate, what to bill, which work to hand over — is answerable at task granularity.
Volume metrics now reward the wrong thing. Commits, lines and deployment counts were weak proxies before; assisted developers commit three to four times more while introducing security findings ten times more, which makes any output-volume metric actively misleading.
And most results will be inconclusive. That is the normal honest outcome for a small team, and it rules out the large effects, which is genuinely useful. The failure is concluding something anyway. DX maintains a useful roundup of developer-productivity research See DX.
One condition runs underneath all of it. The moment anyone suspects the data feeds into judging people, everyone rounds toward the number that looks reasonable — not dishonestly, but rationally — and the measurement stops describing anything. Say explicitly that it is for estimating and pricing, then behave consistently with that.
Cycle Time and What It Hides
Both got faster for reasons that have nothing to do with delivery. What each measures, and the third number that matters more.
Why Measured and Felt Time Disagree
Developers estimated AI made them 20% faster while measurement showed 19% slower. The gap matters more than either number.
When the Numbers Say Nothing
Most small-team data is inconclusive, and that is a result. How to act well without a signal, and what not to conclude instead.
Reading Your Own Data
Small samples, motivated reasoning and survivorship will all tell you a confident story. Five checks before you believe your numbers.
Measuring a Small Team
Most measurement advice assumes an organisation. At two or three people the statistics do not work and the incentives are different.
Tracking People Will Do
Most tracking is abandoned within a fortnight. The failure is never the tool — it is friction, purpose and who reads the numbers.
What to Track, and How Finely
Five fields per task, recorded at the time, beats any dashboard. What to capture, what to skip, and why granularity kills tracking.
Which Tasks Are Fast
Averages are useless when the distribution is bimodal. How to sort your own work into buckets and measure each one separately.