Do You Need Story Points? Estimation for Small Teams
Story points are a relative unit for estimating work, born in Extreme Programming and never part of the Scrum Guide. What they are for, what they cost, the lighter alternatives, and when a small team is better off with them.
7 min read
Most small teams do not need story points. Agile story points are a relative unit for how big a piece of work is, added up after each iteration into a velocity that is then used to forecast. They help when a team plans in fixed iterations, keeps its estimating habits steady and genuinely needs a forecast of how much will fit. Scrum does not require them, and the person often credited with inventing them has said he would rather teams did without. For a team of two to eight people, splitting work until every item is small and counting what gets finished usually tells you as much, for far less talk.
What story points are
A story point is a number that says how big an item is compared with other items, not how many hours it will take. A team agrees a small, familiar item is a 1 or a 2 and sizes the rest against it, often on a scale with gaps that widen as items grow, such as 1, 2, 3, 5, 8, 13. The gaps are deliberate: nobody can tell a 9 from a 10, but most people can tell a 3 from an 8.
The Agile Alliance glossary (opens in a new tab) describes points as possibly the most widespread of the units agile teams prefer to the old man-day or man-hour. Their appeal is that they are dimensionless. Add up the points a team finished in an iteration and you get its velocity; divide the points left by the velocity and you get a rough number of iterations. Nobody has to argue about whether a day of work means a day on the calendar.
Where they came from
Usage varies and there is no single origin, but the best first-hand account comes from Ron Jeffries, one of the creators of Extreme Programming. In Story Points Revisited (opens in a new tab), written in 2019, he explains that XP first estimated stories in time, then in “ideal days”: how long a pair would take if left alone. They multiplied ideal days by a load factor, which tended to be about three, to get real time. Stakeholders were confused about why a day’s work kept taking three days, so the team started calling ideal days “points”.
His summary is often quoted: he may have invented story points, and if he did, he is sorry now. In the same article he says he deplores their misuse, thinks using them to predict when work will be done is at best a weak idea, thinks tracking estimates against actuals is at best wasteful, and thinks comparing teams on velocity is harmful. Other units came and went around the same time: the Agile Alliance lists “Gummi Bears” and “Nebulous Units of Time” among them.
What the Scrum Guide says about estimation
Nothing about points. The 2020 Scrum Guide (opens in a new tab) says Product Backlog items gain attributes such as a description, an order and a size, that the attributes vary with the domain of work, and that the Developers who will do the work are responsible for the sizing. For Sprint Planning it says that the more the Developers know about their past performance, their upcoming capacity and their Definition of Done, the more confident they will be in their Sprint forecasts. It does not mention story points, velocity or planning poker.
So a Scrum team can size in points, in T-shirt sizes, in days or by splitting until everything is roughly the same size. What it cannot do is hand sizing to someone who will not do the work. If your team adopted points because it believed Scrum required them, that belief is worth revisiting.
What points cost
Estimating is not free, and the costs are easy to miss because they are spread across every refinement session:
- Time. Every item gets discussed until the numbers converge. On a small team that can be a quarter of the refinement session.
- False precision. A 5 feels like a measurement. It is an opinion, and the Agile Alliance’s own entry on relative estimation (opens in a new tab) calls the evidence that relative estimates beat absolute ones tentative at best.
- Drift. Once velocity is watched, points inflate. A 3 last year is a 5 this year, and the chart goes up while the output does not.
- Comparison. Two teams’ points are not the same unit, but someone will put them on one slide.
- The wrong question. Arguing over a 5 or an 8 is time not spent asking whether the item is worth doing, or whether it could be smaller.
Four alternatives
T-shirt sizes
XS, S, M, L, XL. The same relative idea as points, but the letters resist arithmetic, so nobody is tempted to add them up and call it a forecast. Give each size a rough meaning, such as under an hour, an afternoon, a day or two, about a week, and treat the largest size as a signal to split rather than a size to plan with.
Count items instead of points
If items are roughly similar in size, the number finished per week, the throughput, forecasts about as well as velocity. Twelve items finished in each of the last four weeks and forty items left gives you a rough answer without a single estimate. It works best alongside the next idea.
Split until every item is small
This is what Jeffries prefers: slice larger stories into smaller ones, each as valuable as possible and ideally less than a day of work. The only judgement left is “small enough” or “not small enough”, which is quicker and more honest than choosing between a 5 and an 8. Once items are small and similar, counting them is a forecast.
No estimates at all
Some teams stop estimating entirely and pick the next most valuable thing each time. In a 2013 piece on the NoEstimates movement (opens in a new tab), Jeffries argued that if you deliver small chunks of work incrementally there is little need to estimate stories, and that predictions can come from measuring how long things actually take. This suits product teams shipping continuously. It suits a fixed-price bid badly.
When points are worth it
Points earn their cost in a narrower set of cases than their popularity suggests:
- The team works in fixed iterations and has to commit to a scope each time, so it needs a consistent sense of how much fits.
- Items genuinely vary a lot in size and cannot sensibly be split, so counting them would mislead.
- Someone outside the team needs a release forecast, and the team has several iterations of stable history to base it on.
- The estimating conversation itself is valuable: a wide spread between a 2 and a 13 reveals that people understand the item differently.
That last point is the best argument for any estimation. The number matters less than the moment two people discover they were imagining different work. You can get that conversation with T-shirt sizes too.
A simple rule for a small team
- Size each item near the top of the backlog with a T-shirt letter, in the team’s regular refinement session. Take seconds, not minutes.
- Anything bigger than a few days gets split before it is started.
- Count what you finish each week. After a month you have a throughput you can forecast from.
- Revisit only if the forecasts are consistently wrong in a way that costs you, and then ask whether splitting further fixes it before adding points.
If you plan in sprints, bring that throughput to the sprint planning meeting as your evidence of past performance. It serves the same purpose velocity would.
Sizing on a fenbs board
fenbs has no story points, no velocity and no sprints. Each task has an optional size from XS to XL: XS is under an hour, S an afternoon, M a day or two, L about a week, and XL means too big, split it. XL shows loudly on the card for that reason. Sizes are letters on the board and are not converted to points anywhere you can see.
- An AI assistant connected over MCP can set a size when it creates or updates a task, with
sizeonfenbs_create_itemorfenbs_update_item, and should propose a split when it finds an XL. - Priority is separate, from 1 (most urgent) to 10, so size and urgency never share one number.
- Throughput needs no feature: History records every move between lanes with who made it and when, so you can count how many tasks reached Completed in a week by hand.
Related
Where sizing happens: backlog refinement. How small a work item should be: user story vs task. Breaking large work down: AI agent task decomposition. Forecasting from flow instead of velocity: what is Scrumban.