Metrics and Key Performance Indicators (KPIs)
Metrics and key performance indicators (KPIs) are the technique that builds a monitoring-and-evaluation system around a solution: choosing the indicators, setting their targets and organising the collection, analysis and reporting of the data that feeds them. An indicator measures the degree of progress toward a goal; a metric is its level read at a given moment; a key performance indicator (KPI) is an indicator tied to a strategic objective. The technique's value lies in designing the system that produces these numbers period after period: a dashboard is drawn in an hour, a system that fills it faithfully for three years is designed.
Goal
The technique builds the system through which an organisation knows whether its solution is meeting its objective and by how much it falls short. BABOK calls this a monitoring-and-evaluation system. Monitoring is the continuous collection of data comparing what was achieved against the expected result. Evaluation is the systematic, objective assessment of the solution over time, to establish its effectiveness and identify where to improve it. Choosing good indicators is not enough: monitoring and evaluation require a durable mechanism that feeds them every period.
The deliverable is this complete mechanism. It comprises the indicators tied to the objectives, their targets and above all the procedures that keep the whole thing running: how the data is collected, how it is analysed, how it is reported and against which baseline progress is measured. It is this system that makes alignment visible, linking goals to objectives, objectives to solutions and solutions to tasks and resources. An indicator that is well chosen but that nothing collects regularly measures nothing at all.
Usage
When to use it
- A deployed solution to evaluate over time: set up the monitoring that will compare the result obtained against the result expected.
- A strategic objective to make steerable: link goals, objectives, activities and resources through indicators fed at regular intervals.
- A change case to support: establish the quantified baseline of current performance before committing the investment.
- A service agreement or programme to oversee: turn the agreed thresholds into targets that are collected and reported.
- A process improvement under way: instrument the before and after to establish the real gain.
- Recurring reporting to set up: fix the procedure once (templates, recipients, frequency) rather than gathering numbers for every committee.
When not to use it
- No strategic objective stated: nothing distinguishes a key indicator from any other number, run the strategy analysis first that sets goals and expected outcomes.
- Measurement meant to rate individuals: it will be gamed at the expense of unmeasured activity, reserve the indicator for steering the system and hand individual appraisal to an objectives review.
- Data source absent or prohibitively costly: the system will produce invented numbers, restrict the scope to indicators whose data exists or build a proxy before committing to primary collection.
An indicator tied to a strategy
The word key ties the indicator to a strategic objective. BABOK reserves the term KPI for an indicator that measures progress toward a strategic objective, and this requirement of attachment rules out most of the numbers tracked day to day. An appointment book's occupancy rate is a useful operational indicator, it becomes a KPI only when tied to a strategic objective it is a lever for. This is the work Kaplan and Norton's Balanced Scorecard formalised: linking indicators to financial, customer, process and learning objectives, so that no tracked number floats free of an intention. A measurement system that has not made this attachment measures activity, not performance.
The four pieces of the system
BABOK states that setting up a monitoring-and-evaluation system takes four things: a collection procedure, an analysis procedure, a reporting procedure and the collection of a baseline. These four pieces are the heart of the technique, and it is their design that decides whether the system holds.
The collection procedure
It fixes six points: the units of analysis (what is being counted, a patient, a course of treatment, an appointment), the sampling procedure when not everything is measured, the collection instruments, the frequency and the responsibility for collection, that is, who does it. The collection method is the first cost item of the whole system: data already recorded by a system is almost free, a survey or direct observation has to be paid for. This trade-off is made before an indicator is retained, because an excellent indicator that is impossible to collect is worth nothing.
The analysis procedure
It specifies the processing applied to the raw data and, just as important, the consumer of the analysis, who often has a strong interest in how it is conducted. The same missed-appointment data is analysed differently depending on whether the aim is to size the front desk or to spread a workload across therapists, and the choice of method is never neutral. Naming the consumer up front stops the analysis being redone at every committee according to the mood of the day.
The reporting procedure
It covers the report templates, the recipients, the frequency and the means of communication. The report typically compares the baseline, the current value and the target, with the gap in absolute and relative terms. Two BABOK principles govern its form: the trend weighs more than the absolute value, and the visual presentation beats the raw table, especially when paired with text that explains what the numbers show.
The baseline
The baseline is the data taken just before or at the start of the period to be measured. It serves twice: to know recent performance and to measure progress from that point. It is collected, analysed and reported for each indicator, without exception, failing which a current number stays unreadable because nothing says where it starts from.
Three qualities of the data
BABOK judges the quality of indicators and their metrics on three criteria. Reliability is the stability and consistency of collection across time and space: an access time counted in working days one year and calendar days the next is not reliable. Validity is the degree to which the data directly measures the intended performance: a high occupancy rate is a sign of good health only if it tracks viability rather than merely how full the appointment book is. Timeliness is the fit between the frequency and latency of the data and the need for a decision: accurate data that arrives after the decision is of no use.
Choosing the indicator and setting the target
A good indicator, in BABOK, has six qualities: clear (a single reading), relevant (appropriate to the concern), economical (reasonable cost), adequate (a sufficient basis to assess), quantifiable (independently verifiable) and trustworthy and credible (grounded in evidence). They often pull against each other, and the trade-off sometimes produces a proxy, a substitute used when the factor cannot be measured directly or when collecting it directly is too heavy. Absent a satisfaction survey, a practice takes the rate of treatment courses carried through to the end as a proxy for the clinical outcome. The proxy is an accepted compromise, and its weakness must be known to whoever reads the number.
The target metric is the objective to reach within a set period. Setting one calls for a clear baseline, a sense of the resources that can be mobilised and an awareness of the political stakes. A well-formed target is specific, measurable and time-bound, which Doran summed up as SMART. It takes the form of a point, a threshold or a range, the last being useful while the indicator is new and its behaviour poorly understood. A leading indicator measures a cause upstream (access time signals the margin to come), a lagging indicator measures an achieved result (the margin of a closed month): a useful system mixes the two, failing which it produces a report card rather than a lever.
Why a measurement system fails
The limitations BABOK notes lie in the system.
It collects too much
Gathering more data than needed costs in collection, analysis and reporting and distracts the team. BABOK notes that the failing is a particular risk on agile projects. The test is simple: if an indicator leads to no decision that will actually be taken, it leaves the system.
It does not loop back
A bureaucratic programme fails by collecting a great deal and producing few useful reports in time. Those who capture the data must get feedback on what their work changes in the result, otherwise the quality of capture degrades and the system is emptied of its substance.
It is turned against people
When a metric is used to rate individuals, they optimise the measure at the expense of unmeasured activities: this is Goodhart's law, as Strathern phrased it, a measure that becomes a target ceases to be a good measure. A practice that rated its therapists on occupancy rate alone would see them favour short sessions at the expense of the heavy cases that drive the outcome. The safeguard lies in the design: measure the collective outcome, cross indicators that check one another and separate steering the system from appraising people.
AI considerations
AI mainly helps to sketch the system. From a stated objective, a model proposes a first list of indicators, their formulas, a skeleton collection procedure and a report template, which one starts from as a draft to be refuted. On the data flow it is solid: detecting anomalies and breaks in a series, spotting a drift in the trend and drafting the qualitative commentary that accompanies a chart. The diff from one report to the next, mechanical and exhaustive, is exactly the vigilance that human review loses over time.
Three decisions stay human. Whether an indicator is key is a judgement about its link to strategy, often political. The target depends on the real baseline and the resources: a model proposes a plausible number, not an attainable one. And AI set loose on optimising a metric is Goodhart's law at machine speed, a system that learns to maximise a number will find the shortcut that inflates it without serving the objective, which leaves the analyst as the guardian of validity. Data sensitivity calls for caution: salaries or health data within the meaning of the revFADP do not go into a general-purpose model.
Examples
The case is a physiotherapy practice in the canton of Vaud, three therapists, whose strategic objective for the year is to offer fast access and treatments carried through to the end while staying viable. The system keeps five indicators. Three are tied to that objective, so KPIs, and two are leading operational indicators that steer them. The definition table is the first piece of the system: it fixes for each indicator its formula, its source, its cadence and its collection responsibility, before any value is even read.
| Indicator | Formula | Source | Cadence | Baseline | Target (12 months) | Type | Key? |
|---|---|---|---|---|---|---|---|
| Access time to 1st appointment | median working days between request and 1st appointment | scheduling software | monthly | 9 d | ≤ 5 d | leading | KPI |
| Treatment-course completion rate | courses completed per prescription / courses prescribed | patient record | quarterly | 68% | 80% | lagging | KPI (outcome proxy) |
| Operating margin per therapist (FTE) | (revenue − costs) / number of therapist FTEs | accounting | monthly | CHF 4'200 | CHF 6'000 | lagging | KPI |
| Appointment-book occupancy rate | slots honoured / slots opened | scheduling software | weekly | 72% | 85% | leading | indicator |
| No-show rate | missed appointments / scheduled appointments | scheduling software | weekly | 11.2% | ≤ 5% | leading | indicator |
Reporting is the piece the system produces every period. It lines up three columns, baseline, current value and target, and quantifies the gap twice, in absolute and relative terms, to make legible the size of the movement and the share of the road still to go.
| Key indicator | Baseline (January) | Current (June) | Target (December) | Gap vs baseline | Gap vs target |
|---|---|---|---|---|---|
| Access time to 1st appointment | 9 d | 6 d | 5 d | −3 d (−33%) | +1 d, 75% of the way covered |
| Treatment-course completion rate | 68% | 74% | 80% | +6 pts (+9%) | −6 pts, 50% of the way covered |
| Margin per therapist (FTE) | CHF 4'200 | CHF 5'100 | CHF 6'000 | +CHF 900 (+21%) | −CHF 900, 50% of the way covered |
Cost
| Phase | Level | Justification |
|---|---|---|
| Preparation | Medium | Designing the four procedures and linking the indicators to the objectives means bringing together the strategy stakeholders and settling who collects what. |
| Execution | Variable | The collection procedure drives the cost: a source already recorded is almost free, primary collection (survey, observation) tips it toward high. |
| Documentation | High | The system is sustained over time: procedures to keep, a report to produce every period, baseline and targets to revise. It is the item that decides its survival beyond the first year. |
Tools
The spreadsheet honestly houses a small system, and that is where the majority of real dashboards live. Visualisation and reporting tools (Power BI, Tableau, Grafana, Metabase) carry the reporting procedure as soon as several sources have to be consolidated, trends traced and a threshold alerted on. The source systems (scheduling, ERP, accounting, ticketing, CRM) hold the raw data, and a data warehouse becomes useful when an indicator crosses several of them.
For the strategic piece, balanced-scorecard or OKR tools give substance to the goal-objective-indicator chain that makes an indicator a KPI. For data quality, a statistical control chart (SPC) serves reliability and trend reading directly, separating normal variation from a real break. The choice of tool follows the design of the system: a fine tool wired to fuzzy procedures produces a dashboard no one links to a decision.
Sources
- IIBA, A Guide to the Business Analysis Body of Knowledge (BABOK Guide) v3, §10.28 Metrics and Key Performance Indicators (KPIs): the indicator / metric / KPI / reporting distinction, the four pieces of the monitoring-and-evaluation system, the six qualities of an indicator, the three qualities of the data, its strengths and its limitations.
- ISO 22400-1:2014, Key performance indicators for manufacturing operations management, Part 1: the conceptual framework and terminology that formalise the notion of a KPI.
- ISO 22400-2:2014, Part 2, Definitions and descriptions: the definition of a KPI by its formula, its elements, its time behaviour and its unit, the reference for structuring an indicator.
- Doran, G. T. (1981), There's a S.M.A.R.T. Way to Write Management's Goals and Objectives: the origin of the SMART criteria behind the specific, measurable, time-bound target metric.
- Kaplan, R. S. & Norton, D. P. (1992), The Balanced Scorecard, Measures That Drive Performance: the link between indicators and strategic objectives, what makes an indicator a key one.
- Strathern, M. (1997), Improving Ratings: Audit in the British University System: the modern formulation of Goodhart's law, the anchor for the limitation on gaming and sub-optimisation of measures.

