Your Training Partner
Techniques Toolbox
The Delphi cycle in six steps: questionnaire, anonymous estimates, aggregation into a median and interquartile range, controlled feedback without the identities, revised estimates, then a stopping rule that starts a new round or delivers the converged estimate.

Delphi Method

The Delphi method turns the judgement of a group of experts into a single, defensible estimate through several rounds of anonymous questionnaires. After each round a coordinator returns the group statistic to the panel, a median and its interquartile range, together with the arguments behind the extreme estimates; each expert then revises their figure in the light of what they did not know. Because the estimates are formed without anyone knowing who proposed what, rank, politics and the bandwagon effect carry no weight in the result, and the spread narrows round after round toward a consensus. BABOK places it among the estimation methods, alongside bottom-up, top-down and parametric estimation, rough order of magnitude and PERT.

Goal

The Delphi method draws from a panel of experts an estimate that uses all the available knowledge without being corrupted by the group dynamics that spoil an open estimation meeting. It serves the decisions that call for a figure no calculation can produce: no comparable history to calibrate a model against, no parametric basis from which to derive a value, the only source of information being the knowledge of a handful of experts who do not agree. Asking a single expert yields a fragile figure, dependent on their own vantage point. Putting the experts in one room yields something else, a figure coloured by everyone's rank, by the first amount said out loud and by the comfort of going along with the group.

The decision it supports is a commitment made under uncertainty and without hard data: a cost, an effort in person-days, a duration, a quantity to forecast, on an initiative novel enough or contested enough that an off-the-cuff figure will not survive a steering committee. Delphi organises that judgement. It gathers the estimates separately, reports back to the panel what the group produced without revealing identities and lets the figures converge over several rounds. It produces a figure accompanied by its spread and by the record of the disagreements that shaped it.

The deliverable is a converged group estimate, expressed as a central value, the median of the final round's estimates and as a measure of spread, the interquartile range, which says how far the panel agreed. To this is added the trail: the statistic of each round and the justifications for the extreme positions, which show why the group moved and make the estimate auditable. Delphi is one of the methods in the estimation family, the one reached for when expert judgement is the only raw material and a single expert will not do.

Usage

When to use it

  • Experts scattered or remote: written, asynchronous rounds where a workshop cannot bring everyone together.
  • Rank or politics would distort open debate: anonymity lets a junior contradict a senior without paying for it.
  • High anchoring or bandwagon risk: first-round estimates form before anyone has seen a group figure.
  • No historical or parametric basis: expert judgement is the only input, and a single expert is too fragile.
  • High-stakes or contested estimate: a novel initiative or a disputed budget justifies the cost of several rounds.

When not to use it

  • Open confrontation is workable: a facilitated workshop or a planning poker session reaches consensus faster and more cheaply.
  • Cheap or urgent decision: a rough order of magnitude (ROM) or a single PERT pass will do.
  • No real expert pool or a panel with one blind spot: anchor on data through parametric estimation or an analogous estimate.

Description

The four features that define the technique

Delphi is a repeated procedure resting on four features that Rowe and Wright identify as the core common to all its variants: anonymity, iteration, controlled feedback of the group response and statistical aggregation of the individual answers. Remove any one of the four and the nature of the exercise changes.

Anonymity is the reason the technique exists. Experts respond separately and no one but the coordinator knows who put forward which figure. This is what neutralises authority, timidity and the bandwagon effect: an estimate is defended by its argument. Iteration arranges several successive rounds, because an expert who discovers an argument they did not know has the right to change their mind, and that right is what a single meeting refuses. Controlled feedback is what the coordinator returns between two rounds: the group statistic and the reasons given for the extreme estimates, never the identity of their authors. Statistical aggregation, finally, summarises a round's answers into a central value and a spread, so that the group sees where it stands without any one voice dominating the report.

The round cycle

One begins by defining precisely what is being estimated and in what unit, because a panel answering slightly different questions never converges. Then comes the first round: each expert delivers their estimate independently, without having seen the others' or any reference figure, which guarantees that the starting point is not already anchored. The coordinator then computes the median and the interquartile range of the round and collects the justifications for the highest and the lowest estimates. They return this controlled feedback to the panel, the median-and-spread pair together with the arguments, without ever saying who said what. In the next round, each expert revises their figure in the light of those arguments or holds their position with a justification. This repeats until a stopping rule fixed in advance: a spread threshold reached or an agreed number of rounds, typically two or three. Convergence is read from the interquartile range tightening from one round to the next. The panel usually numbers five to about twenty experts, chosen more for the diversity of their knowledge than for their number: a disagreement between distinct expertises is what makes the method work.

1Questionnairedefine the estimand2Anonymousestimatesno reference value3Aggregationmedian + IQR4Controlledfeedbacknever the identities5Revisedestimateseach re-assessesSpread below threshold?or rounds reachedyesConverged estimatemedian + IQR (last round)no · new round
The Delphi cycle: anonymous estimates are aggregated into a median and an interquartile range, controlled feedback returns the statistic and the arguments from the extreme positions, and each round tightens the dispersion until the stopping rule.

The group statistic: median and interquartile range

Delphi imposes no calculation formula: the figure comes from the experts. What is fixed is the way a round is summarised. The canonical statistic is the median, not the mean. The reason lies in the technique being designed to surface the extreme estimates: an expert who sees a risk the others miss will propose a figure far off to one side, a valuable signal. A mean lets itself be pulled by that isolated figure and shifts the group summary toward a point almost no one defends; the median resists that extreme while letting it live on in the debate. Spread is measured by the interquartile range, the distance between the first and the third quartile, that is, the extent of the middle half of the answers. An interquartile range that shrinks over the rounds is the operational definition of convergence, and it says more than a median, which can stay stable while the group tightens around it.

What makes the exercise fail

The first pitfall is the gravest because it mimics success. Forced consensus consists of shrinking the interquartile range by pressing the extreme opinions to fall into line. What one then gets is a central tendency with no real agreement: convergence is a signal, never a target to manufacture. A coordinator who leans on the holdouts produces a tight and false figure and destroys the information the rounds were meant to gather.

The second is panel homogeneity. A group that thinks alike converges quickly, and it converges on its shared error: the speed of tightening is reassuring only if the expertises brought together are disparate. The third is coordinator bias. Since the coordinator alone chooses which justifications are returned to the group, they steer the panel by what they highlight or pass over in silence. Feedback must be faithful, including for the arguments that unsettle the median, or the method merely propagates the facilitator's opinion. The fourth is attrition: over several asynchronous rounds, experts drop out, and the aggregate of the last rounds rests on a panel that is depleted and sometimes biased by those who remain. The fifth is the absence of a stopping rule. Without a spread threshold or an agreed number of rounds set in advance, the exercise drags on at the coordinator's patience or stops on an arbitrary round, and the decision to close becomes a source of bias in itself.

The formal procedure and its accidental double

The full Delphi protocol, its numbered rounds, its coordinator and its written stopping rule, rarely runs as such outside foresight exercises and institutional expert panels. What one meets far more often is a degraded form that settles in by accident. In an organisation where teams work in silos, several people put a figure on the same thing each on their own, without talking, and the anonymity that gives the method its value is realised without anyone having intended it. The result has the look of a Delphi first round, with its independent, unanchored estimates, and it stops there. The aggregation into a median and an interquartile range, the controlled feedback that circulates the arguments of the extreme positions and the rounds that tighten the spread all remain undone. One is left with a cloud of figures never brought together, mistaken for an estimate.

The practical consequence comes down to two steps. Recognising this involuntary Delphi for what it is prevents the belief that a shared figure exists when all one has is parallel opinions. Completing it costs little, since the hard part, obtaining independent estimates, is already done: what remains is to summarise the round by its median and spread, to return the reasons for the extreme positions to the group without the identities and to run a second round. The method begins the moment people who were ignoring each other start to confront their figures under a rule.

AI considerations

The coordinator's mechanical load is where an assistant is most useful. Aggregating a round's estimates, computing the median and the interquartile range, tracking the tightening of the spread round after round, all of this is calculation without judgement that a model, or even a spreadsheet, performs without failing. Beyond the arithmetic, a language model helps handle the free-text justifications: grouping them by theme, faithfully summarising the arguments of the extreme positions, preparing legible feedback for the next round and flagging when the spread stops decreasing, that is, when the stopping rule is near. This is a real gain, because careful feedback is the time-consuming part of running the exercise and the part on which the quality of the next round depends.

What a model cannot do lies in the very nature of the technique. A language model is not a panellist. An estimate it produced has no field experience behind it, and it reintroduces the bias Delphi exists to remove: it reflects the consensus of its training data, a massive and invisible anchoring on what is already written. Choosing the panel, reading why a dissenting expert holds their position and above all telling honest convergence from forced convergence remain human judgements. Finally, the individual estimates and justifications are personal data under the revised Swiss Data Protection Act (nFADP) and can be politically sensitive: the anonymity that gives the method all its value must survive whatever tool aggregates, which rules out feeding named responses into a public service whose data handling is not under control.

Examples

A health insurer in French-speaking Switzerland estimates the effort, in person-days, of building a new online portal for submitting claims. Five seasoned analysts and architects sit in different offices, which makes a synchronous workshop awkward, and the project has no historical analogue to calibrate a model against, which rules out parametric estimation. Three anonymous rounds are run; the experts are labelled E1 to E5, and the median and the interquartile range are computed each round under a single convention held constant, the first quartile taken as the second ranked value and the third as the fourth.

Delphi method

Online claims portal, effort in person-days

PanellistRound 1Round 2Round 3
E1607590
E2758595
E39095100
E4120110105
E5200140120
Median9095100
Interquartile range452510
Three anonymous rounds: the median settles around ninety to a hundred person-days while the interquartile range collapses from 45 to 25 then to 10. It is the spread that converges.

Everything turns on the contrast between the median and the interquartile range. The median barely moves, from ninety to a hundred person-days, and an observer watching only it would think the first round nearly right. What changes is the interquartile range, which falls from 45 to 10: in the first round the panel was scattered, from sixty to two hundred person-days, and by the third the middle half of the answers sits within ten days. The engine of the tightening is the feedback. E5, in justifying a very high figure, exposed a data-migration risk the others had not costed; the argument, returned anonymously, pulled E1, E2 and E3 upward, while E5, seeing that the risk was serious but bounded, revised downward. The convergence was not imposed; it came from the information that circulated without anyone's rank weighing on it. The figure committed to is the final round's median, a hundred person-days, carried with its residual spread.

Visualisations

The technique produces two objects, and each calls for its own medium. The first is the convergence table: made of rows and columns, it is the deliverable itself, and its median and interquartile-range rows are what a reviewer recomputes to check that the convergence is real. It stays in HTML, selectable and recomputable, rather than frozen into an image.

The second is the procedure itself, which is better seen drawn than described. The drawing carries the round cycle: the questionnaire, the anonymous individual estimates, the aggregation into a median and interquartile range, the controlled feedback that returns the statistic and the arguments of the extreme positions, then the revised estimates, with a loop back as long as the stopping rule is not met. It shows at a glance what a list of features only names: that anonymity and feedback are two faces of one loop, repeated until convergence.

Cost

PhaseLevelRationale
PreparationMediumThe real cost is human: recruiting a panel whose expertises are disparate, defining without ambiguity what is being estimated and in what unit, writing the questionnaire and fixing the stopping rule before the first round. A vague estimand or a homogeneous panel dooms everything that follows.
ExecutionHighSeveral asynchronous rounds spread over days to weeks, with, each round, the aggregation of the statistic, the collection and faithful summary of the extreme justifications, the drafting of the feedback and the reminders to contain attrition. This is the long and most demanding part of the technique.
DocumentationLow to mediumThe statistic of each round and the arguments that moved it. The care goes into the traceability of the convergence, which makes the estimate auditable, and into preserving anonymity in the records kept.

Tooling

A simple online form coupled with a spreadsheet is enough for a panel of a few experts: the form gathers the estimates separately and preserves anonymity toward the panel, the spreadsheet computes the median and the interquartile range of each round in a few cells and replays the calculation the moment a value changes. For most estimates, nothing more is needed.

Dedicated Delphi platforms automate the cycle when the panel is large or the rounds are repeated: they handle sending the questionnaires, anonymisation, automatic aggregation and the presentation of the feedback, and some run a real-time Delphi in which the feedback updates continuously, which shortens the schedule at the price of finer control to exercise over how the statistic is presented. This extra tooling is justified when the size of the panel or the frequency of the exercises makes manual running unmanageable; below that, it adds a licence without contributing anything a form and a spreadsheet do not already do. Whatever the tool, the non-negotiable point is that anonymity and the fidelity of the feedback never depend on the platform's convenience.

Sources

  • N. C. Dalkey and O. Helmer, "An Experimental Application of the Delphi Method to the Use of Experts", Management Science, 9(3), 1963, pp. 458-467: the first open publication of the method developed at the RAND Corporation, which sets out obtaining a reliable consensus through a series of questionnaires interspersed with controlled opinion feedback.
  • G. Rowe and G. Wright, "Expert Opinions in Forecasting: The Role of the Delphi Technique", in J. S. Armstrong (ed.), Principles of Forecasting, Kluwer, 2001, pp. 125-144: the methodological consolidation that identifies the four constitutive features (anonymity, iteration, controlled feedback, statistical aggregation), the size and diversity of the panel and the content of the feedback between rounds.
  • IIBA, A Guide to the Business Analysis Body of Knowledge (BABOK Guide) v3, §10.19 Estimation: the descriptive placement of the Delphi method among the expert-judgement estimation methods. BABOK situates the technique within that family and does not prescribe its procedure, which rests on the work of Dalkey and Helmer and of Rowe and Wright.
Decision Trees
All techniques
Descriptive and Inferential Statistics