Status Droid how long software work actually takes

What We Cannot Measure

A site arguing that you should measure rather than trust your impressions owes you the boundary of that argument.

Here it is: most of what determines whether software work goes well cannot be measured by a small team, and some of it cannot be measured at all. Pretending otherwise is how measurement turns into theatre. A related workplace concept is covered in learn more, which is useful context for this argument.

The six

Whether the thing was worth building. The largest question and entirely outside the data. A team can be perfectly calibrated on estimates and spend a year building something nobody needed. No metric in this site's whole scheme detects that.

Whether the approach was right. You can measure how long the approach took. You cannot measure the alternative you did not take, because it does not exist. Every architecture decision is an uncontrolled experiment with a sample size of one.

Long-term consequences of today's shortcuts. A decision that saves two days now and costs three weeks in eighteen months registers as a saving. The cost appears later, attributed to something else, in a different quarter.

Quality that never fails. Code that quietly does not break generates no incident, no rework flag, no data at all. The best work is invisible to measurement, which means any metric will systematically undervalue it. The Verge provides mainstream technology reporting that helps frame tool and workplace shifts See The Verge.

Judgement. Knowing which task is which, when to stop reviewing, when to say the requirement is wrong. The scarce input, and there is no counter for it.

Whether people want to keep doing this. Sustainability of practice, which determines everything over a five-year horizon and appears in no dashboard until someone leaves.

Why this is not an argument against measuring

The case for measurement was never that numbers capture everything. It was narrower.

Perception of speed is demonstrably miscalibrated, by around 39 points in the best evidence available. Measurement corrects one specific, documented failure of intuition. It does not replace intuition, and on the six above intuition remains the only instrument there is.

So the honest position is a division of labour. Count what is countable and use it for the decisions it serves — estimating and pricing. Use judgement for everything else, and know which you are doing.

The failure mode is not measuring too little. It is letting the measurable crowd out the important, which happens by default because one produces charts and the other does not.

How the crowding-out actually happens

Not by decision. By attention.

You start tracking cycle time. Cycle time becomes visible. Visible things get discussed, and discussed things get optimised. Six months later the team is choosing smaller tasks, which improves the number, and nobody chose that.

Nothing in this sequence involves anyone acting badly. It is what happens when one part of a system is instrumented and the rest is not.

The defence is to say out loud, regularly, what the numbers do not cover. Which is most of the reason this page exists.

What to do about the unmeasurable

Ask directly. Was this worth building? Would you take this approach again? The answers are unreliable and they are better than nothing, and nobody asks.

Keep a decision log. Not metrics — one line per significant choice, and what you expected. Six months later you can check. This is the closest thing to measurement available for architecture decisions.

Watch for the things that do not appear. Rising rework with stable throughput. A quiet person. A codebase nobody volunteers to work in. These are signals with no number attached.

Trust the feeling where it is reliable. It is accurate about interest, interruption, trouble and tiredness — just not about quantity. Ask it the questions it can answer.

The honest summary of this whole site

Measure the small number of things that are countable and that feed decisions you actually make. Do not mistake that for knowing whether the work is good.

The counting exists to correct one known bias, not to replace judgement, and a site that argued otherwise would have overclaimed exactly the way it accuses everyone else of overclaiming.

The short version