Relative Estimation
Relative estimation gives a backlog item a size by comparing it with items that already have one, on a scale of unitless numbers, the story points. An item at 5 is taken to be about five thirds of an item at 3. The number carries no duration. The Agile Extension to the BABOK Guide sets out estimation in an agile setting: the sizing is redone iteration after iteration and grows more accurate as the team learns its own capacity and the nature of the work. The points completed within an iteration add up to the team's velocity. Points belong to the team that set them and do not carry across to another team.
Goal
Relative estimation sizes and orders a backlog without putting hours or francs on it from the outset. The Agile Extension gives it three uses: working out the cost and effort of a body of work, setting the priorities of the initiative, committing to a schedule. The guide places estimates outside the delivered solution: they serve the internal decisions of the team and its sponsors.
What the technique produces is a backlog in which every item carries a size on one shared scale, together with the velocity recorded over the previous iterations. The two together answer the question a steering committee asks: how many iterations for this scope.
Place in the estimation family
Estimation in the general sense covers the methods that produce a quantity in a real unit, a duration, a cost or an effort, even where that comes as a range. Relative estimation produces a rank on a dimensionless scale. The quantity appears only afterwards, when the team's velocity turns a point total into a number of iterations, then that number into francs at the observed cost of one iteration. The Agile Extension builds on the general technique and describes its adaptation to an agile setting.
An effort in person-days compares across teams, adds up across projects and goes into a contract. A unitless rank serves one purpose, the ordering of a backlog by one given team, and most of the errors come from asking it for the other three.
Usage
When to use it
- A backlog to order before a commitment: comparing the items with one another is enough to settle the delivery sequence.
- A stable team that delivers in iterations: the observed velocity turns the points into a number of iterations.
- Items still loosely defined: the uncertainty goes into the size without any need to break the item down into tasks.
- Choosing between the initiatives of a portfolio: placing the scale of the candidates before committing to a detailed costing.
- The opening of an iteration: each item receives its size in session, with no individual preparation beforehand.
When not to use it
- A contractual commitment in hours or in francs: put numbers on it directly, by bottom-up or parametric estimation.
- A team reshuffled at every iteration: velocity has no baseline to work from, so break the work down and price it in units of time.
Description
What the number measures
The Agile Extension bases the size of a user story on five factors, which the session puts to each item in turn. Knowledge: what information does the team hold about this item. Experience: has it already done this or something close to it. Complexity: how hard is the implementation. Size: how much work does the item ask for. Uncertainty: what variables and unknown factors can throw it off.
One of these factors on its own is enough to put an item at the top of the scale. Two days of work where nobody knows how the neighbouring system behaves is rated above a week of work the team has already been through three times. The number bundles an amount of work together with a degree of ignorance.
The scale and the reference story
The usual scale is a modified Fibonacci sequence, rounded at the top: 1, 2, 3, 5, 8, 13, 20. The gaps widen as the numbers climb, which rules out an argument between 12 and 13 over an item almost nothing is known about. Two neighbouring values at the bottom of the scale cover work whose difference the team can feel; two neighbouring values at the top cover work it cannot tell apart. T-shirt sizes (XS, S, M, L, XL) apply the same ordinal principle to labels.
The team picks a reference story: an item already delivered, small, understood by everyone, which it rates a 2 or a 3. Every other size is set against that one. An item bigger than the largest card in the deck calls for a split, which the session flags before the item enters an iteration.
Getting the scale started
The Agile Extension offers three starting points to a team with no history yet. The first is the rough order of magnitude: size roughly, adjust over the iterations that follow. The second starts from a given set of resources and an iteration of fixed length, then lets the first iteration show what fits inside it. The third has the team estimate a sample of stories of differing sizes in time, then extrapolates what an iteration can absorb.
How accurate these three starting points are matters little: the Agile Extension holds that the first estimates are coarser than those produced as delivery approaches, and the technique draws its value from the correction each iteration brings.
Planning poker
Planning poker sizes the items in session, with the whole team, usually during the planning workshop. The team goes through a backlog item until everyone understands it the same way. Each member then picks a card on the point scale in private, and all the cards are turned over at once. If the values sit within one step of each other, the team keeps one and moves to the next item. If the spread is wide, it discusses the knowledge one person has and the others do not, the knowledge that explains why the same item looks like a 3 to one and a 13 to another. The session aims at agreement on each item.
The whole mechanism rests on the simultaneous reveal: it stops the first value spoken from fixing the ones that follow. James Grenning formalised the technique in 2002; Mike Cohn brought it into agile practice in 2005.
Magic estimation
Magic estimation, which the Agile Extension calls Silent Sizing in §7.13.4, handles a whole backlog by moving cards around, without speaking. The team prepares a deck of cards, one item per card. It also needs a wall or a table to lay the cards out in a line. Each person in turn takes a card and places it on the line, the smallest items at one end, the largest at the other; a member may move a card someone has already placed. The round repeats until every card is on the line. The team then looks for the breaks, the places where the size gap between two neighbouring cards widens, then forms groups. Only at that point does each group receive a value from the scale, and each card takes the value of its group.
The relative position is settled before any number exists. The price is the silence: the knowledge a discussion would have drawn out stays with whoever holds it.
Before grouping, position only
After grouping, the value appears
Velocity and the conversion into a schedule
Velocity is the total of the points completed in an iteration. Over several iterations it measures the team's throughput, on which its next commitments rest. It is read as a range: a team that has delivered 34, 41, 38 and 45 points announces a range of 34 to 45 points. A single average, 39.5, would give the committee a precision that four readings do not carry.
The point total of the scope, divided by the low velocity and then by the high one, gives a number of iterations as a range. The cost follows from multiplying that number by the fully loaded cost of one iteration, the figure cost accounting already holds. The conversion therefore always runs through velocity. A price per point bypasses it and cuts the only link that tied the points to a real quantity.
What throws the estimate off
The fixed conversion rate
Declaring that a point is worth four hours puts back the unit the technique leaves out. The ratio frozen that way shifts at every change in the team's make-up. The team then rates the items in hours in its head and divides to find the card to play. The five factors disappear, and uncertainty with them. What is left is an estimate in hours dressed up as points.
Comparison between teams
The numbers belong to the team that set them. A 5 in one team has no bearing on a 5 in another, and the Agile Extension lists comparison between teams among the limitations of the technique, for the confusion it creates among stakeholders. Compared velocity becomes a performance indicator, and a team measured on its velocity raises it within one iteration by rating higher. The number climbs and the work delivered stays where it was.
Anchoring in the session
An architect who announces a value before the cards are turned over takes from the session what it was producing: judgements formed independently, then compared. The same effect arrives unannounced through the round of the table where cards are shown from left to right. So the order in which cards are shown is settled in advance.
Unannounced recalibration
A team recalibrates its scale over time: what it rated a 3 six months ago it rates a 2 today. The move is healthy, but it makes the old velocity incomparable with the new one. Recalibrating without saying so produces a velocity curve that seems to drop for no reason and a committee discussion about a fall in performance that never happened. The answer is to date the recalibration and to start again from the new range.
The estimate read as a date
The Agile Extension flags two misreadings by stakeholders outside the team. The first takes the relative estimate for a firm deadline. In the second, attention settles on the estimate, which is an output, when the result being sought is the value delivered. The answer is to present the conversion: a range of iterations and the cost that follows from it, revised at every iteration. A committee shown a point total will always want to know what a point is worth.
AI considerations
Before the session, a language model works on the backlog and on the history held in the tool. It picks out the items whose description does not allow their size to be judged, those that overlap an item already delivered whose past size is the point of comparison, those that exceed the largest card in the deck and for which it proposes a split. On the history, it computes the velocity range, isolates the atypical iterations and flags a drift in the scale by comparing the average size of items from one quarter to the next.
The number itself stays with the team. A model asked for the size of an item produces a plausible value drawn from other projects, whereas the five factors all bear on what this team knows, has already done and does not know. The value therefore arrives calibrated on a scale foreign to the history the team built its own from. The risk is greater still when the model is put inside the loop of the session: a proposal displayed before the cards are turned over is an anchor like any other. A backlog also holds customer segments, amounts and regulatory constraints that have no business leaving the company, and these are stripped out before anything goes to a public model.
Examples
In a Swiss retail bank, the digital banking product team sizes the items of the mortgage simulation module of its mobile app, ahead of the steering committee's budget review. The deck runs from 1 to 20 and the reference story, rated 2, is the addition of a field to the simulation form, delivered two iterations earlier.
| Backlog item | Size | The factor that drove the number |
|---|---|---|
| PDF export of the simulation result | 2 | Size: existing template, one screen, no new rule. |
| Biometric login through the existing e-banking app | 3 | Experience: the team has integrated the same OAuth mechanism twice already. |
| Affordability check against the bank's internal lending rules | 5 | Complexity: seven calculation rules, each with its own thresholds and exceptions. |
| Real-time mortgage rate feed from a new external provider | 13 | Uncertainty: nobody has ever integrated this provider's interface. |
| Migration of the saved simulations from the old portal | 20 | Knowledge: the original storage format is documented nowhere. |
The rate feed started out split between 5 and 13. The discussion brought out that the provider was new to everyone, and the team settled on 13 for that reason alone, with no change in the volume of work expected. The migration of the simulations sits at the ceiling of the deck for a related reason: nothing the team has delivered resembles it. The accuracy of the technique rests on that resemblance.
Over the previous four iterations, the team completed 38, 44, 36 and 42 points, a range of 36 to 44 points per two-week iteration. The scope put to the committee totals 210 points, which gives five to six iterations. At the fully loaded cost of one iteration, CHF 45'000 at the current headcount, the committee receives a range of CHF 225'000 to CHF 270'000 and the date that goes with it. At no point does it receive a price per point.
Visualisations
The drawing that teaches the technique is the magic estimation line. A first row shows the item cards laid down one after another, smallest to largest, with no number in sight. The same row, below, is cut at the gap breaks into four groups, and each group then receives its value from the scale.
The table of the sized backlog is the second figure: one item per row, its size and the factor that explains it.
Cost
| Phase | Level | Justification |
|---|---|---|
| Preparation | Low | A refined backlog and a deck of cards, physical or online. The reference story is chosen once and for all. |
| Execution | Medium | One to two hours with the full team, at every iteration. Magic estimation covers the whole backlog in a single pass, with no item-by-item discussion. |
| Documentation | Low | The size is an attribute of the item in the backlog tool, and velocity is computed there on its own. |
Tooling
Backlog management tools (Jira, Azure DevOps, GitLab and their equivalents) hold the point field on the item and compute velocity from what has been completed. They are also the source of the history the technique lives on: without a reliable reading per iteration, velocity is an opinion.
Simultaneous voting tools (planning poker apps, collaborative whiteboard templates) make the simultaneous reveal workable for a distributed team, which a video call with no voting device does not: someone there always calls out a number before the others.
The physical deck remains the fastest option for a co-located team, and the wall with its cards is the only setting where magic estimation runs at its natural speed. Remotely, a collaborative whiteboard reproduces the line and the groups, at the price of slower handling.
The spreadsheet is enough to hold the velocity range and the conversion into iterations and francs, a few rows to show the steering committee.
Sources
- IIBA, Agile Extension to the BABOK Guide, §7.13 Relative Estimation: the purpose of the technique, the progressive character of estimation in an agile setting, the three contributions to stakeholders, the five factors behind the size of a story, the Fibonacci scale, the definition of velocity, the three ways of starting a scale, planning poker and Silent Sizing, together with the strengths and limitations stated there, among them the dependence of accuracy on how far new stories resemble those already delivered, comparison between teams, the estimate read as a firm deadline and the attention paid to the output rather than to the outcome.
- IIBA, Agile Extension to the BABOK Guide, §4.7.1, §5.7.1 and §6.7.1, the techniques by planning horizon: the use of the technique at the Strategy horizon to place the value and the resources of the initiatives in a portfolio and at the Initiative and Delivery horizons to decide which features to deliver and in what order.
- IIBA, A Guide to the Business Analysis Body of Knowledge (BABOK Guide) v3, §10.19 Estimation: the general estimation technique, which the Agile Extension builds on and whose adaptation to an agile setting it describes.
- Mike Cohn, Agile Estimating and Planning, Prentice Hall, 2005: the practice reference that spread story points, the reference story, the Fibonacci deck rounded at the top and the use of velocity as the single route of conversion to a schedule.
- Mike Cohn, Agile Estimating: How Teams Estimate with Story Points, Mountain Goat Software: the statement of the distinction between relative effort and absolute duration, together with the team-specific character of the point scale.
- Agile Alliance, Planning Poker glossary entry: how the technique runs and its formalisation by James Grenning in 2002, then its spread by Mike Cohn in 2005.

