Status Droid how long software work actually takes

Measuring

The case for measuring is narrow and it is strong: perception of speed is demonstrably miscalibrated, by around 39 percentage points in the best evidence available, in the direction that flatters the tool.

Measurement corrects that one documented failure. It does not tell you whether the work is good, and nothing in this section pretends otherwise. For a practical implementation reference alongside the measurement argument here, see learn more.

What to record is smaller than you think. Five fields per task — timestamps, category, whether an assistant was used, the estimate written down before starting, and a flag for rework. Fifteen seconds. Anything costlier gets abandoned inside a fortnight, and inconsistent data is worse than none because you will trust it anyway.

Granularity is the trap. METR labelled 143 hours of screen recordings to get second-level resolution. That is a research budget. Every decision a small team actually makes — how to estimate, what to bill, which work to hand over — is answerable at task granularity.

Volume metrics now reward the wrong thing. Commits, lines and deployment counts were weak proxies before; assisted developers commit three to four times more while introducing security findings ten times more, which makes any output-volume metric actively misleading.

And most results will be inconclusive. That is the normal honest outcome for a small team, and it rules out the large effects, which is genuinely useful. The failure is concluding something anyway. DX maintains a useful roundup of developer-productivity research See DX.

One condition runs underneath all of it. The moment anyone suspects the data feeds into judging people, everyone rounds toward the number that looks reasonable — not dishonestly, but rationally — and the measurement stops describing anything. Say explicitly that it is for estimating and pricing, then behave consistently with that.