Software Estimation: Why It Goes Wrong and What Works
Software estimates go wrong for predictable reasons: single numbers that hide uncertainty, optimism about the work in front of you, and estimates mistaken for promises. The methods that do better are ranges, comparison with past work, small pieces and coarse sizes.
7 min read
Software estimation goes wrong for the same few reasons on almost every team. People give one number when they honestly mean a range. They picture the work in front of them and forget how long similar work took last time. And the estimate, once spoken, gets treated as a promise. What works is the reverse of each: give a range with a stated confidence, check it against how long comparable work actually took, split anything big until it is small enough to see, and size in coarse buckets instead of hours. For a lot of day-to-day work, a size and a rule that big things get split is all the estimation a small team needs.
Why software estimates go wrong
- Optimism about the case in front of you. Psychologists call it the planning fallacy: people forecast from the details of the plan and underweight how things usually go.
- Single numbers. “Two weeks” sounds precise, but it hides whether you mean “probably two” or “two if nothing surprises us.”
- Estimates confused with targets and commitments. An estimate is what you think will happen; a target is what someone wants; a commitment is what you agree to deliver. Mixing them turns every estimate into a negotiation.
- Work that is too big to see. A three-month item has unknowns nobody has found yet, so any number for it is mostly a guess.
- Missing work. The estimate covers writing the code and forgets review, testing, deploy, data migration and the questions that come back.
The Agile Alliance’s glossary entry on estimation (opens in a new tab) names three of the same mistakes: estimates always carry uncertainty, so “point” estimates “are generally considered inadequate”; estimates are not commitments, and blaming a developer for taking three days on a two-day estimate is counterproductive; and an estimate reflects what was known at the time, so it should always be permissible to update it.
Give ranges, not single numbers
A range with a confidence says more than a point estimate and costs no extra effort: “three to six weeks, and I am about 80% sure it lands in that window.” The width of the range is information. A narrow range says the work is familiar; a wide one says there are unknowns worth finding before anyone commits.
- Ask for the range first and the most likely value second, so the range is thought through on its own rather than drawn as a margin around one number.
- Say what would move you to the top of the range: the vendor API is worse than the docs, the data needs cleaning, the reviewer is out.
- Keep the confidence honest. If six of your last ten “80% ranges” were missed, they were really 40% ranges.
- Narrow the range by doing the risky part first, not by thinking harder about it.
Large programs formalize the same idea. NASA’s Cost Estimating Handbook (opens in a new tab) includes an appendix on joint cost and schedule confidence level analysis, which expresses a plan as a probability of finishing within a given cost and date rather than as a single figure. A small team does not need the math, but it can borrow the habit of saying how sure it is.
Reference class forecasting: compare with past work
The strongest fix for optimism is to stop estimating from the inside. Dan Lovallo and Daniel Kahneman, in their Harvard Business Review article Delusions of Success (opens in a new tab), call this taking the “outside view”: ignore the details of your project at first, look at how a class of similar projects actually turned out, and place yours within that spread. Their steps, adapted to software:
- Pick a reference class: the last ten or so pieces of work that were genuinely similar, such as “new integration with a third-party API.”
- Look at how long they actually took, the typical case and the extremes, from your records rather than memory.
- Place the new work in that spread using what you know about it: easier API, same team, messier data.
- Ask how well your past gut calls matched reality, and pull the estimate back toward the typical case as far as they missed.
Bent Flyvbjerg developed this into reference class forecasting (opens in a new tab) for large infrastructure projects, basing forecasts on the actual performance of comparable projects so as to bypass optimism bias. The software version only needs one thing most teams already have: a record of when work started and when it finished. That is why a board with a history of moves is worth more to estimation than any estimating technique.
Split the work until it is small
Estimates for small items are better than estimates for large ones, not because people are better at small numbers but because a small item has fewer places for surprises to hide. The practical rule: if you cannot say what the first day of work looks like, the item is too big to estimate. Split it by outcome, not by layer: “customer can export invoices as CSV” and “export includes line items” rather than “build the backend” and “build the frontend.”
Splitting is also the step AI assistants help with most. Given a clear goal, an assistant can propose a breakdown for a person to prune; AI agent task decomposition covers how to ask for one.
Size in coarse buckets
Once items are small, you rarely need hours. A coarse size such as XS, S, M, L, XL is quick to agree on and honest about its precision. Give each size a rough meaning, and make the largest a signal to split rather than a size to plan with. Whether you also need story points and velocity is its own question, answered in do you need story points; for most small teams the answer is no.
And sometimes no estimates at all
Some teams stop estimating individual items entirely and forecast by counting how many small items they finish each week. Ron Jeffries, writing about the NoEstimates movement (opens in a new tab), argued that teams delivering small increments have little need to estimate stories, and that predictions can come from measuring how long things actually take. It fits a product team shipping continuously. It fits a fixed-price proposal badly, where someone still needs a range before work begins.
Estimating work an AI agent will do
When a coding agent writes the first draft, the typing part of a task often shrinks, and the parts around it do not: writing a clear task, reviewing the result, testing it and fixing what the agent misunderstood. Estimate those explicitly. A task that is well specified and easy to verify stays close to its size; a vague one tends to come back more than once. How to write a task for an AI agent is the cheapest way to keep that second kind rare.
Estimating on a fenbs board
fenbs has no story points, no velocity, no sprints and no due dates. Each task has an optional size, and the meanings are fixed so everyone reads them the same way: XS is under an hour, S an afternoon, M a day or two, L about a week, and XL means too big, split it. XL shows loudly on the card for that reason.
- Priority is a separate 1 to 10 scale, 1 the most urgent, so how big and how urgent never share a number.
- History records every move between lanes with who made it and when, so your reference class is on the board: look at when similar tasks moved to In Progress and when they reached Completed.
- Put the range and what would push it to the top in the task’s plan, and rewrite the plan as you learn. The note stays the problem.
- An AI assistant connected over MCP can set
sizeonfenbs_create_itemorfenbs_update_item, and should propose a split when it finds an XL. - If your team agrees “anything L or bigger gets split before it starts,” record it on the Decisions and rules page as a rule, so every connected assistant follows it too.
Related
Points, velocity and the lighter alternatives: do you need story points. Where sizing happens: backlog refinement. Measuring how long work really takes: cycle time vs lead time. Paying back the shortcuts estimates leave behind: technical debt.