Benchmarking
Benchmarking compares an organisation's performance or practice against an external reference point: a leading performer, a peer cohort or a published standard. It locates the gap between current practice and the best demonstrated practice, then converts that gap into an improvement target. An internal number tells you whether you are improving; it does not tell you how far you stand from what is achievable. Benchmarking supplies the external reference that judges the number and the practice that explains how the leader achieves it.
Goal
Benchmarking supplies the external reference that turns an absolute number into a judged one. An organisation can track its own performance for years without knowing whether that performance is good: an internal trend answers "are we improving?", it does not answer "how far are we from what is achievable?". A cost of CHF 14.20 to process one invoice means nothing until a peer cohort sits at CHF 9.80 and the best-in-class performer at CHF 5.60.
The decision it supports comes down to two questions: where to direct the improvement effort and how far to aim. The gap sizes the stake, it says whether the subject warrants a project and the leader's practice suggests the means, what the best do differently. Benchmarking also informs make-or-buy decisions and the verification of compliance against a standard, where the reference is a published norm.
The deliverable is a gap analysis: the measures chosen for the organisation, set against the reference cohort and the best-in-class performer, the gap computed in absolute and relative terms and, wherever the study also covers practice, the enabling practices that produce the leader's result. To this are added an improvement target that the gap justifies and the project proposal to reach it. Benchmarking is one of the two halves of the composite technique that the BABOK guide files under benchmarking and market analysis, alongside market analysis, which studies a market, its segments and its players.
Usage
When to use it
- An absolute performance with no reference: an internal number does not say whether it is good, an external reference judges it.
- Framing an improvement objective: the gap to the best sizes the prize and says whether a project is warranted.
- A multi-site organisation: one site already outperforms the others, its practice transfers internally at least cost.
- A make-or-buy decision: compare your performance against a provider or a peer before deciding.
- A compliance check: measure the gap to a published standard.
- The search for the practice behind the number: understand how the leader gets its result.
When not to use it
- No comparable external reference: for a genuinely new product or process, establish the baseline through a current-state capability analysis.
- A breakthrough goal: to aim for a durable advantage that outruns the leader, open the options through an ideation workshop.
- A decision too small for the cost: for a reversible operational tweak, a root-cause analysis on your own process is far cheaper.
Description
Two ways to compare
Two independent cuts structure a benchmark. The first names whom you compare against, the second what you compare. They cross, and conflating them is the first framing error.
Whom, that is the classic taxonomy inherited from Robert Camp. Internal benchmarking compares one unit, site or team against another inside the same organisation. It is the fastest and cheapest, the data is owned and the practice transfers with the least loss of context, but the ceiling is your own best unit: you learn nothing the organisation does not already do somewhere. Competitive benchmarking compares against a direct competitor in the same market. It is the one you need to know where you stand against the firms you actually lose deals to, and it is also the hardest, because a competitor does not share its numbers and the legal and ethical limits bite hardest here. Functional benchmarking compares against an organisation that performs the same function in a different sector, a logistics process against another industry's logistics: the process is comparable even though the businesses are not, and the pool of partners willing to exchange is far wider than in the competitive case. Generic benchmarking pushes the same logic to its limit: comparing against a best-in-class performer of a generic process, billing, order fulfilment, complaint handling, regardless of sector. It is the highest ceiling available, at the cost of the most abstraction: the further the source, the more adaptation the practice needs. The last three types together form external benchmarking; internal is the fourth.
What you compare crosses that first cut. Performance benchmarking confronts quantitative measures, indicators: cost, cycle time, defect rate. It tells you how far behind you are. Practice benchmarking confronts how the work is done, the qualitative "how": people, process, technology. It tells you why the gap exists and what to adapt. The mature study does both, in that order: performance data locates the lag, the practice study explains it and shows what to change. Performance alone tells you the leader is faster, it does not tell you how it got there, and a gap without its cause does not turn into action.
Running the technique
The original model, the one Camp formalised at Xerox, has five phases: planning, analysis, integration, action and maturity. The cycle a practitioner actually runs comes down to five moves, and naming Camp's phases behind them lets you recognise the reference model.
- Plan: decide what to compare and select the cohort
Choose a subject, a process or an output that matters and that you can measure, define the exact measure and its unit of comparison, then pick the reference partners. The subject should be one where a gap would genuinely justify a project. - Collect: gather our data and theirs
Internal data first, because you cannot read a gap you cannot see on your own side. Then the partners': published studies and specialised databases, surveys, a request for information on capabilities, site visits to best-in-class organisations. Record the basis of every number: definition, scope, costs included and period. - Normalise: make the comparison fair before reading the gap
This is the step amateurs skip and the one that decides whether the result is real. Two "costs per invoice" are not comparable if one includes the fully loaded cost, employer AVS and LPP contributions, systems and overhead, and the other only salaries; if one counts invoices and the other invoice lines; if scopes, volumes or accounting periods differ. Restate every measure onto the same definition, the same scope, the same cost basis and the same period before subtracting. - Analyse: read the gap and study the practice behind it
Compute the gap, in absolute and relative terms, then, for anything you intend to act on, study why the leader is ahead: the enabling practice behind the number. - Set the target and adapt the practice
Set an improvement target that the gap justifies, often the cohort median in the first cycle and the best-in-class extreme after several, then adapt the leader's practice to your own context. These are Camp's integration and action phases: fold the target into plans, implement and re-benchmark periodically, the maturity phase, because the reference moves.
The pitfalls that make the gap false
The first and commonest is to compare non-comparable quantities. Different measure definitions, scopes, cost bases, volumes or periods make the gap fictional. This is exactly what normalisation prevents, and it is why normalisation is not negotiable: a gap computed on unaligned measures lies, and it lies all the more for looking precise.
The second is to copy a practice without its context. A best-in-class performer's practice is tuned to its scale, culture, regulation and technology. Transplanted whole into a different context, it under-performs or fails. Benchmarking transfers ideas to adapt.
The third is metric fixation: optimising the benchmarked number at the expense of what it stood for, driving the cost per call down by pushing customers to a channel they hate. You benchmark a small balanced set of measures.
The fourth is data availability and quality. Benchmarking is time-consuming, and the organisation does not always have the expertise to conduct the analysis. Competitor data is often unavailable or unreliable, and a partner's self-reported data is uneven. A gap computed on bad data is worse than no gap, because it steers a project in a false direction with the assurance of a number.
The fifth is legal and ethical, and it bears above all on competitive benchmarking. Gathering data on a competitor must never slide into industrial espionage or collusion. The recognised guardrail is the Benchmarking Code of Conduct (APQC and the Global Benchmarking Network): legality, reciprocity of exchange, confidentiality, use limited to improvement and no discussion of prices or market allocation that would breach competition law. It is the rule of the game for any exchange between partners.
The last is a limit of the technique itself: benchmarking does not produce innovation. It assesses solutions shown to work elsewhere, so at best it catches up with the current leader, it does not overtake it and does not build a durable competitive advantage. Framing it as catching up avoids asking of it what it cannot give.
AI considerations
The first useful use is collecting public performance data at scale. An assistant sweeps annual reports, sector studies, regulatory filings and industry databases for comparators and their reported indicators, far faster than a manual search. The second, the most rewarding, is normalising and reconciling measures: flagging that two numbers rest on different cost bases or scopes, proposing a common definition, restating the numbers onto it. This is mechanical and error-prone by hand, exactly where a tool earns its place, under review. The third is spotting candidate comparators, those organisations performing a comparable function in another sector that you would not have thought to search, the functional and generic pool. The fourth is drafting the gap analysis and the study write-up from the normalised table.
Where AI must not replace judgement, three points are sensitive. Context transfer first: deciding whether a leader's practice will survive being moved into a different scale, culture, regulation and technology is a judgement about two specific organisations. An AI that recommends copying a practice has no model of the receiving context and will encourage exactly the copy-without-context pitfall. Data provenance and sensitivity next: a model will happily assemble data on a competitor without ever asking how it was obtained. The legality and ethics of a source, the Code of Conduct, are the analyst's responsibility, and a partner's data received under an exchange agreement is never fed into an external model. Trust in harvested numbers last: a number an assistant extracted carries no guarantee of its basis, and every number that enters the comparison must be checked for its provenance and normalised by a human before the gap is read.
Examples
An industrial SME in the canton of Vaud benchmarks the cost of processing one supplier invoice in its accounts-payable process, against a cohort of Swiss peers, over 60'000 invoices a year. The concept the comparison makes visible is that one and the same number, our cost, is judged against two references.
Gap analysis
Cost of processing one supplier invoice
| Measure | Our organisation | Swiss cohort median | Best-in-class |
|---|---|---|---|
| Cost of processing one invoice | CHF 14.20 | CHF 9.80first-cycle target | CHF 5.60 |
| Absolute gap | CHF 4.40 | CHF 8.60 | |
| Relative gap | +45 % | 2.5× | |
| Annualised gap (60'000 invoices) | CHF 264'000 | CHF 516'000 |
The target for the first improvement cycle is the cohort median at CHF 9.80, worth CHF 264'000 a year: best-in-class lights up the maximum possible gain but is not aimed at in a single step. A method note reads inside the artifact itself: the CHF 14.20 is fully loaded, and the cohort numbers were restated onto that same basis before the gap was read, without which a peer counting only salaries would look artificially cheap.
Visualisations
The gap analysis is made of rows and columns, and the deliverable is the table itself. Carrying it as HTML rather than an image has practical consequences: the gaps recompute when a measure changes, it stays legible at zoom and to a screen reader and its target column stands out to the eye. An image of a grid freezes numbers that the slightest revision of a measure makes wrong.
Two conventions make the reading immediate. Our value and the two references read on a single row, because the gap to each is precisely what the technique teaches. And the actionable reference is marked, the cohort median, so that the first-cycle target is seen without commentary, distinct from best-in-class, which marks the extreme.
The gap itself is better drawn than tabulated, because it is a distance. A horizontal-bar comparison of the three values, with the gap bracket drawn between our bar and the median and the target marker set on that median, carries the central idea in one image: the gap is the distance to a chosen reference, and the near reference is the target you act on.
Cost
| Phase | Level | Justification |
|---|---|---|
| Preparation | High | The real cost is here: choosing the right subject and the right measure, then above all selecting and recruiting the reference partners. Lining up a competitive or functional cohort takes real effort. |
| Execution | Medium | Collection, surveys, request for information, site visits and normalisation take time, but stay bounded once the cohort is set. |
| Documentation | Medium | The gap analysis and the project proposal are a genuine write-up, but produced once per study. |
Tooling
The spreadsheet is the honest and sufficient choice for the comparison and the gap computation: the measures line up in columns, the absolute and relative gap is a formula and the annualisation a multiplication. A shared document gathers the practice study in parallel, the "how" that accompanies the "how much".
For the reference data, specialised sources make the difference on cohort quality. The APQC Open Standards Benchmarking programme and its process classification framework supply standardised process measures, comparable by construction. Sector-federation surveys and published sector and regulatory data feed the peer data; in Switzerland, the federal and cantonal statistics of the Federal Statistical Office add to it. Beyond that, subscriptions to benchmarking databases and firms that run managed studies open cohorts a single organisation would not reach, and business-intelligence tooling normalises and visualises the gap once the data is assembled.
None of these tools is a blank cheque on data exchange. The Benchmarking Code of Conduct governs the use of all of these tools: it sets what a partner may ask, exchange and reuse, as well as what competition law forbids raising.
Sources
- IIBA, A Guide to the Business Analysis Body of Knowledge (BABOK Guide) v3, §10.4 Benchmarking and Market Analysis: the definition, the elements of the approach (identifying the leading enterprises, competitors included, surveys, request for information, site visits, determining the gaps, project proposal), the mention of competitors, government and industry associations as sources of best practice and the recognised limitations, including that it is time-consuming and cannot produce an innovative solution. Descriptive claims only.
- Robert C. Camp, Benchmarking: The Search for Industry Best Practices that Lead to Superior Performance, ASQC Quality Press, 1989: the founding work of formal benchmarking, developed at Xerox, and the five-phase process model (planning, analysis, integration, action, maturity). Primary source of the technique, of which Camp is the inventor.
- APQC (American Productivity & Quality Center), What Is Benchmarking? and What Are the Four Types of Benchmarking?: the taxonomy of types (internal/external crossed with performance/practice) and the method, from the recognised authority of the field.
- APQC and the Global Benchmarking Network, Benchmarking Code of Conduct: the legal and ethical guardrail for gathering and exchanging benchmarking data (legality, exchange, confidentiality, use, no discussion of prices or markets).

