Your Training Partner
Techniques Toolbox
Sampling distribution in four parts. On the left, the population of transactions of a retailer in French-speaking Switzerland as a cloud of points, crossed by the line of the population mean, labelled unknown and carrying no figure. In the centre, four samples of four hundred transactions cut out of the cloud, each with its own mean. On the right, the four means placed on an axis in francs, spread on either side of the population line. Below them, the confidence interval band built on the first sample, wide enough to cover the other three. At the bottom, across the full width, twenty stacked intervals: nineteen cross the line of the true value, one misses it.

Descriptive and Inferential Statistics

Descriptive statistics summarise the data set in hand: where the values concentrate, how far they spread, what shape their distribution takes. Inferential statistics extend that summary to a population nobody has measured in full, from a sample drawn out of it. The deliverable changes in kind: the answer is an estimate carrying an interval and a confidence level.

Goal

Descriptive and inferential statistics are the measures that summarise a data set and extend that summary to the population the data is drawn from. The technique applies when a decision calls for a figure nobody has measured in full. A retailer's average basket, the satisfaction of its customers, the handling time of a claim at an insurer: each is estimated on a sample.

The IIBA guide locates the difficulty in the interpretation: quantify the data, then read it in its business context, before it feeds a decision.

The deliverable comes in two pieces. Descriptive statistics give a measure of location, a measure of spread and the shape of the distribution, on data whose type allows it. The estimate gives the value sought, its confidence interval and the sampling conditions that make the inference legitimate.

Usage

When to use it

  • A population out of reach of a census: a few hundred observations drawn at random estimate a mean to within a few per cent; the adoption rate, the willingness to pay or the satisfaction level goes into a business case with its interval.
  • Comparing two periods or two variants: two means can be compared only once their spread is known.
  • An indicator to define or to monitor by sampling: defect rate, response time, volume of complaints, whose measures of location and spread are chosen before a numerical target is set.
  • An order of magnitude disputed between stakeholders: the calculation does not change with who reads it, a strength the IIBA guide points out.

When not to use it

  • No data measured on the factors at work: an expert judgement structured by the Delphi method yields more than a falsely precise figure.
  • Several uncertainties combined by a formula: business simulation carries the ranges through to the result.

Description

Two distinct questions

 Descriptive statisticsInferential statistics
QuestionWhat does the data in front of me say?What can I claim about the population I have not measured?
ObjectThe observed data set, whatever its originA sample drawn from a defined population
ConditionsA data type compatible with the measure computedRandom selection, independent observations, sufficient size
ResultLocation, spread, shape of the distributionAn estimate, its interval and its confidence level
What it does not sayNothing about what was not observedNothing about the cause of what is estimated
The two halves of the technique, row by row: the object, the conditions and the deliverable all differ. The last row marks the boundary with hypothesis testing and exploratory data analysis.

Two neighbouring subjects stay outside the scope. Deciding whether a gap observed between two groups comes from sampling variation belongs to hypothesis formulation and testing, with its null hypothesis, its significance level and its rejection decision. Looking at the raw data before summarising it, searching it for outliers and gaps, belongs to exploratory data analysis. The confidence interval and the hypothesis test nevertheless read the same sampling distribution, so a hypothesis placing the mean outside the 95% interval is rejected by the two-sided test at the 5% level.

The data type decides the measure

Continuous data takes any value in a range: revenue, a duration. Discrete data is counted in whole units: a headcount, a number of items ordered. Categorical data names a class with no order: a canton, a product range. When the classes are ordered, the data is ordinal: a school grade from 1 to 6, a satisfaction scale from "very dissatisfied" to "very satisfied".

The IIBA guide stresses the consequence: treating ordinal data as a category erases the order, and an excellent grade then weighs as much as a failing one. A mean computed on a satisfaction scale assumes that the gap between "dissatisfied" and "neutral" equals the gap between "satisfied" and "very satisfied", which nothing guarantees. An average satisfaction of 3.7 on five steps, an average risk level of 2.4, an average priority of 1.8: such numbers populate dashboards and come out of a calculation the data does not permit.

Where the values concentrate

The mean is the sum of the values divided by their count. The median cuts the data into two halves of equal size. The mode is the most frequent value. Quantiles cut the distribution into slices of equal size, percentiles into hundredths and quartiles into quarters.

On a symmetrical bell-shaped distribution, mean, median and mode coincide, as the NIST handbook notes, so the choice does not show. It shows as soon as the distribution leans. The mean follows the extreme values; the median resists them. The mode splits in two when several values are equally frequent, which makes it unwieldy as a single summary. The median is also used to fill missing data, a practice the IIBA guide points out.

How far they spread

Two funds offered by the same bank both show a 3.2% average annual return over ten years. The standard deviation of the first is 1.1%, that of the second 9.4%. The mean is identical and the two investments do not carry the same risk. The IIBA guide makes this its argument for measures of spread.

The variance is the mean of the squared deviations from the mean. The standard deviation is its square root. Both carry the same information and only one of them is communicated: the variance is expressed in squared units, square francs, which the guide flags as an obstacle for business stakeholders. The standard deviation returns to the unit of the mean and is reported next to it.

Skewness says which way the distribution leans; kurtosis says how heavy its tails are, hence how often extreme values occur. These are diagnostics: they decide whether the mean deserves to be reported. Marked right skew, the ordinary case for amounts and durations, puts the mean above the median and above what most observations are worth.

Shape or data typeLocationSpreadWhat to watch
Continuous, symmetrical distributionMeanStandard deviationSubgroups behaving differently, which the pair hides
Continuous, skewed distribution or with extreme valuesMedianInterquartile rangeThe size of the extremes, which the interquartile range ignores by construction
OrdinalMedianDistribution by classThe actual gap between two steps, which is not measured
CategoricalModeCounts per classNo order between the classes: neither median nor mean makes sense
Discrete with a small rangeMedian or modeRange, countsA mean with decimals matches no possible observation
The decision table the reader keeps: which measure of location and which measure of spread are admissible for which shape or data type and what to watch once the pair is chosen.

Population, sample, parameter, statistic

ISO 3534-1 sets the vocabulary of four terms that practice confuses. The population is the set of units the question bears on: the 52'000 transactions of a quarter, the claims of one insurance line. The parameter is the quantity sought on that population, its mean for instance, and it stays unknown as long as the population is not measured in full. The sample is the observed subset. The statistic is the same quantity computed on that sample.

The statistic estimates the parameter and does not equal it. The distinction drives the calculation: the variance of a sample used to estimate that of the population is divided by the sample size minus one, as the NIST handbook gives it; the IIBA guide notes only that the moments about the mean are adjusted.

Why two samples never give the same figure

Drawing a second sample from the same population gives another mean. These means spread around the unknown parameter in a pattern of their own, the sampling distribution. Inference rests entirely on it. Its spread is measured by the standard error, the standard deviation of the sample divided by the square root of the sample size.

Precision improves as the square root of the sample size. Quadrupling the sample halves the standard error: going from 400 to 1'600 observations costs four times as much for a gain of a factor of two. The size of the population does not enter the calculation: estimating a mean over a canton of 300'000 inhabitants or over the whole of Switzerland calls for a sample of comparable size at the same precision.

Population52'000 transactions in the quarterpopulation meanunknown
Four samplesn = 400 per sampleCHF 68.40CHF 66.10CHF 69.30CHF 67.20
The means, plottedCHF 6466687072CHF 65.34CHF 71.46

Twenty intervals, built the same way

each segment is the confidence interval of a new sample

19of 20 contain the population mean
1of 20 misses it
One sample gives one mean; four samples give four. The confidence interval is built on that spread: of twenty intervals computed the same way, about nineteen contain the population mean nobody has measured.

The confidence interval

The confidence interval is an interval computed on the sample by a method whose long-run success rate is known; Jerzy Neyman set out the complete theory in 1937. Its common form for a mean is the sample mean plus or minus a multiple of the standard error, read from Student's t distribution for the confidence level and the sample size. The usual level is 95%.

A 95% confidence interval does not say that there is a 95% chance of the true mean lying inside it. The interval computed on one given sample either contains the true value or does not. The confidence level is a property of the method: were the sample drawn a great many times and the interval built the same way each time, about 95% of those intervals would contain the population mean.

What makes the inference legitimate

Three conditions make the inference defensible. The selection is random, meaning every unit of the population has a known and non-zero probability of being included in the sample. The observations are independent: four answers collected from four people in the same department, after they have talked it over, do not count as four observations. The size is sufficient for the difference the decision has to distinguish, which is computed before collection.

The Federal Statistical Office applies these conditions in the Swiss Labour Force Survey: each quarter, a sample drawn from a register yields estimates published with their confidence intervals. The ILO unemployment rate comes from that survey. The SECO unemployment rate, the one the press quotes every month, counts the people registered with the regional employment centres and rests on no sample.

The IIBA guide puts first among its limitations the obligation to identify and record the assumptions the method makes. Write beside the figure: the population, the sampling frame, the selection method, the sample size, the non-response rate and what may set the non-respondents apart. An estimate delivered without those lines is not auditable.

Expected value, when the outcome is uncertain by nature

When the outcome is uncertain, the expected value is the mean of the possible values weighted by their probability. An insurer expecting motor claims of CHF 2'000, CHF 1'500 and CHF 1'000 with probabilities of 0.5, 0.3 and 0.2 anticipates a charge of CHF 1'650 per policy. No claim will be worth that amount; it is nevertheless the figure that sets the premium.

The IIBA guide flags a framing error here: optimising the wrong quantity. An organisation aiming at a lower staff turnover rate optimises a probability, whereas its outcome depends on the expected value of the business lost: the departure of a key account manager does not weigh the same as that of a role with no customer impact.

Two variables taken together

Covariance and correlation measure how two variables vary together, the second being a normalised version of the first, bounded between -1 and 1. A strong correlation signals a relationship worth examining, without saying which way the causality runs or whether a third factor explains both. Modelling that relationship to predict one variable from the other belongs to regression.

The pitfalls

The interval read as a probability

The steering committee takes away "there is a 95% chance the average basket is between 65 and 71 francs": the sentence costs them next quarter, when the same calculation returns a different interval.

The large biased sample

A high sample size tightens the interval without correcting anything in a flawed selection. Fifty thousand voluntary responses to an online questionnaire estimate the satisfaction of people who answer online questionnaires, with remarkable precision and a bias that nothing in the calculation flags. Four hundred people drawn at random from the customer base are worth more: the interval then describes sampling error alone.

The single figure that outlives its interval

The estimate leaves the analysis with its interval and reaches the committee without it, because a slide holds a number better than a range. The remedy: write the lower and the upper bound in the same cell as the estimate, from the spreadsheet onward, so that no copy can separate them.

AI considerations

The first useful application is the choice of measure: given the structure of the data set, the type of each column and the question asked, a language model proposes the admissible measures of location and spread and flags the ordinal column on which a mean is being computed. The check is mechanical and the eye skips it in a dashboard of forty indicators.

The second is writing the calculation: a few lines of code that produce the descriptive statistics, the standard error and the interval with the right correction and the right distribution, within reach of an analyst who does not program. The same code sizes the sample before collection, a short calculation few teams perform.

A language model asked "what is the average basket in Swiss retail" produces a plausible number without having drawn a sample and without giving an interval, for want of a population behind it. The quality of the selection is a property of the organisation that collected the data. Data sets often contain personal data whose transmission to an external service makes the company liable under the Federal Act on Data Protection: the calculation is done on aggregates or inside a controlled environment.

Examples

A retailer in French-speaking Switzerland estimates the average basket of the past quarter in order to size a promotional campaign. The quarter contains 52'000 transactions and the analysis covers 400 of them, drawn at random from the sales ledger.

ItemValueReading
Population52'000 transactions in the quarterThe parameter sought is the mean of that set
Sample400 transactions, simple random selection from the sales ledgerEqual and known selection probability for every transaction
Sample meanCHF 68.40The point estimate of the average basket
MedianCHF 54.20One transaction in two stays below CHF 54.20
Sample standard deviationCHF 31.20The spread of the baskets around their mean
Skewness+1.4Right-leaning distribution: a few large baskets pull the mean above the median
Standard errorCHF 1.5631.20 divided by the square root of 400, the spread of the sample means
95% confidence intervalCHF 65.34 to CHF 71.4668.40 plus or minus 1.96 times the standard error
Recorded assumptionsComplete sampling frame, no weighting, returns and cancellations excludedWhat the estimate assumes and the number does not show
The complete deliverable, on the example of the retailer in French-speaking Switzerland. The point estimate, CHF 68.40, is only one row out of nine: the interval, the descriptive measures and the recorded assumptions are part of the result.

Distribution of the 400 baskets in the sample

CHF 10 bins

median CHF 54.20 mean CHF 68.40
0306090120150150+Number of baskets
The 400 baskets in the sample, in CHF 10 bins. The distribution leans right: a few large baskets pull the mean to CHF 68.40, above the median of CHF 54.20, which one transaction in two does not reach.

The CHF 14 gap between the mean and the median comes from the skew: a promotional campaign calibrated on CHF 68.40 targets a basket most customers do not reach, and the median is the figure to keep for that sizing. The interval, for its part, bounds what the sample allows anyone to claim about the whole quarter.

Visualisations

Two things are worth drawing. The move from the population to the sample means, then from their spread to the interval: the sequence is spatial and a reader who sees it once stops reading the interval as a decorative error bar. Then the gap between mean and median on a leaning distribution, which a histogram shows at once. The descriptive statistics and the estimate are reported in a table, with their reading alongside.

The IIBA guide counts among the limitations of the technique the difficulty of communicating a statistical analysis without visuals or narrative. Business visualizations deal with that point. An estimate is drawn with its interval. A skewed distribution is shown as a histogram or a box plot.

Cost

PhaseLevelRationale
PreparationMediumDefining the population, obtaining a usable sampling frame and sizing the sample take a few days. Collection itself, when it goes through a survey, exceeds that budget.
ExecutionLowThe descriptive statistics and the interval come out of a spreadsheet or fifteen lines of code in minutes, on data already collected.
DocumentationMediumThe assumptions, the selection method, the non-response rate and the domain of validity are recorded, failing which the estimate is neither auditable nor reproducible next quarter.

Tooling

The spreadsheet covers the everyday need. Excel and LibreOffice Calc compute mean, median, quantiles, standard deviation, skewness and kurtosis as native functions, and Excel's Analysis ToolPak adds the descriptive statistics in one block. One precaution: build the interval by hand. The Analysis ToolPak offers it as an option and the functions CONFIDENCE.NORM and CONFIDENCE.T compute it, but no mean formula displays it on its own.

Python and R take over as soon as the analysis has to be re-run every period or put under version control. The pandas, NumPy and SciPy libraries, like R's base functions, produce the full summary and the intervals in a few lines, and the code file documents the method. Statistical packages (SPSS, Stata, jamovi, JASP) serve teams that do not program and return the interval without being asked.

Business intelligence tools (Power BI, Tableau, Qlik) compute the same measures and pose the opposite risk: they display a mean by default, on any numeric column, without raising the question of the data type or showing an interval. An indicator defined in these tools gains from being reviewed against metrics and KPIs. Survey platforms (Qualtrics, LimeSurvey, SurveyMonkey) handle the selection, the reminders and the margin-of-error calculation at collection time, where the problem is settled.

Sources

Delphi Method
All techniques
Desk Check