Tracking People Will Do
The fields are simple. Getting anyone to fill them in for more than three weeks is the actual problem, and it is a people problem wearing a tooling problem's clothes.
Three reasons tracking dies, in order of how often. For a practical implementation reference alongside the measurement argument here, see this guide.
One: it costs too much
Anything over about fifteen seconds per task gets skipped when you are busy, which is when the interesting data is.
Skipping produces the worst outcome — not missing data, but biased data, because the tasks you skip are systematically the stressful and interesting ones. You end up with a clean record of your calm weeks.
The fixes are boring and they work. Fewer fields. Defaults for everything. No free text. One place, not two. And the entry has to be possible in the moment rather than requiring a context switch, because a context switch is the cost, not the typing.
Two: nobody knows what it is for
"We should track our time" gets compliance for a fortnight and then decays, because there is no visible consequence of doing it. The GitHub Blog regularly covers developer workflows, automation, and engineering practice See GitHub Blog.
State the specific question. We are trying to find out whether our estimates on integration work are systematically low. People will record data for a question they want answered. They will not record data for an archive.
Then show them the answer. Six weeks in, spend ten minutes on what the numbers said, including when they said nothing conclusive. That single meeting does more for adherence than any reminder.
Three: it is being used to judge people
The one that kills tracking permanently.
The moment anyone suspects the data feeds into evaluation, everyone rounds toward the number that looks reasonable. Not dishonesty — a rational response. And once it happens you cannot undo it, because the trust is what you were relying on.
Say the purpose explicitly and then behave consistently with it. Data for estimating and pricing, not for judging individuals. If you cannot make that commitment — because you manage people and the information would be relevant — then be honest about it and accept that the numbers will be softer.
For a solo practitioner this is not a concern, which is why solo tracking usually produces better data than team tracking.
What works in practice
Record at start and stop, not at the end of the day. Reconstruction is the failure mode you are trying to avoid — the whole premise is that recall is unreliable.
One entry per task, not per activity. Finer granularity is where good intentions go to die.
Make the estimate part of starting. Not a separate ceremony. You are about to begin something; write down how long you think it will take. Two seconds, and it is the field with the highest return.
Accept gaps. A record with holes in it is usable. A record you abandoned is not. Nobody should feel they have failed at tracking because they forgot a day.
Review on a schedule, briefly. Monthly, fifteen minutes. Enough to keep it alive, not enough to become a ritual.
What to do when it slips
It will. The habit is fragile for the first two months and it dies quietly.
Do not restart with a better system. The instinct is to conclude the tool was wrong and adopt a new one. It was not the tool. Restarting with more structure makes the friction worse and the second attempt fails faster than the first.
Restart smaller. Two fields for a fortnight. Timestamps and estimate. Add the rest when those are automatic.
Notice which tasks you skipped. That pattern is itself information — usually the ones that were going badly, which are the ones you most wanted to see.
The honest expectation
You will not get complete data. You will get maybe 70% of tasks in a good month, weighted slightly toward the ones that went well.
That is enough. Fifteen tasks per bucket gives you a shape, and a shape drawn from 70% of your work beats an intuition drawn from none of it. Perfect tracking is not the alternative on offer; the alternative is going back to a feeling that is known to be miscalibrated.
The short version
- Over fifteen seconds a task and it gets skipped exactly when the data would be interesting
- Skipping produces biased data, not missing data — a clean record of your calm weeks
- State the specific question it answers, then show people the answer within six weeks
- If anyone suspects it feeds evaluation, everyone rounds toward reasonable and it cannot be undone
- Record at start and stop, one entry per task, and put the estimate in at the moment you begin
- When it slips, restart smaller rather than with a better system — the system was not the problem