Your Training Partner
Techniques Toolbox
The hypothesis test sequence, the guide's seven steps grouped into five stages, cut by a vertical barrier: on the left, what is settled before the data is collected; on the right, what is calculated afterwards.

Hypothesis Formulation and Testing

Hypothesis testing is the procedure that sets a quantified statement about a population against a sample of data and says whether, assuming that statement is true, the sample is improbable enough for the statement to be rejected. Two mutually exclusive propositions are written down. The null hypothesis assumes no effect; the alternative hypothesis carries the belief to be checked. The level of improbability at which the null will fall is then fixed, before the data is looked at. The IIBA guide places the technique in problem analysis and locates the analyst's skill in formulating the statement and in reading the result; the calculation belongs to the data team.

Goal

Hypothesis testing settles a business claim nobody has checked: the complaint rate has fallen, this variant costs less. It turns the claim into a statement that data can contradict, then sets it against a sample under a rule fixed in advance. The IIBA guide gives it that function: converting an intuitive judgement into a verifiable and measurable one, so that decisions are not taken on gut feel alone.

The deliverable comes in two pieces. A written protocol is dated before collection and carries the two hypotheses, the variable measured, the test statistic chosen, the sample size, the significance level and the decision rule. A results note then gives the observed value, the calculated statistic, the decision to reject or fail to reject the null hypothesis and the confidence interval.

The technique sits downstream of exploratory data analysis, which brings the candidate pattern to the surface, and it rests on descriptive and inferential statistics for sampling, estimation and the interval. What it adds is a decision rule fixed before the data.

Usage

When to use it

  • Costly belief drawn from experience: a manager asserts a gap nobody has measured and the decision commits a budget.
  • Decision resting on a sample: a survey, a pilot or a panel that has to hold for the whole population.
  • Comparison of two groups or two periods: before and after, canton against canton, where the observed gap may be noise.
  • Result that will be challenged: a supervisory authority, an audit committee, a partner; a protocol predating collection makes the conclusion defensible.

When not to use it

  • Whole population available: descriptive statistics give the gap without a test.
  • Question of cause: the test shows that a gap is real without explaining it; a randomised experiment such as A/B testing establishes the link.
  • No candidate hypothesis: nothing to refute yet; exploratory data analysis supplies the lead.

Description

Two statements that exclude each other

The null hypothesis, written H0, is the default position: no effect, equality with the reference value. The alternative hypothesis, written H1, is its negation and carries the belief to be checked. The test examines only one of the two: it measures how improbable the observed data would be if H0 were true. It rejects H0 when that improbability crosses a threshold fixed in advance.

The direction of H1 governs the shape of the test. A two-tailed test takes any departure from the reference value in either direction and splits the risk between the two ends of the distribution. A one-tailed test takes one direction only and concentrates the risk on one side, which makes it better able to detect a gap in that direction. The direction comes out of the business question and is chosen before collection: switching from a two-tailed test to a one-tailed test once the sign of the gap is known doubles the announced risk.

The threshold is fixed before the data is seen

The significance level, written α, is the risk of rejecting a true null hypothesis. NIST defines it as a value that whoever runs the test sets for themselves, in advance. Professional practice adopts 5%, sometimes 1% where a wrongful rejection carries heavy consequences. The decision rule turns that threshold into a mechanical comparison: reject H0 if the test statistic falls beyond the critical value, on the side H1 points to.

Once the data is in view, any threshold can be justified after the fact and the observation window shortens until the gap appears. The protocol, written and dated before collection, is what separates a result from a demonstration arranged afterwards.

The two errors and what they cost

Decision takenH0 is true (no effect)H0 is false (real effect)
Reject H0Type I error, of probability α: action is taken on an effect that does not existCorrect decision
Fail to reject H0Correct decisionType II error, of probability β: a real effect is let through

The two risks are not managed the same way. α is set freely, since it is the threshold itself. β depends on the real gap between the population and the null hypothesis, which nobody knows, but also on the sample size, the spread and α. It is lowered by increasing the sample or by seeking to detect only a larger gap. 1 − β is called the power of the test, the probability of detecting an effect that exists.

Setting α is a business decision. α is lowered where acting wrongly is expensive and hard to undo, as with a national roll-out or a contract renegotiation. Where missing a real effect costs more, as with a pool of savings left aside year after year, α stays at 5% and the larger sample is paid for to gain power. Costing the two errors in francs before the test is launched makes that choice negotiable in the meeting.

The procedure

The IIBA guide breaks the sequence into seven steps.

State the hypothesis. "The programme reduces costs" cannot be tested. "The average annual claims cost of members enrolled in the programme is below CHF 4'800" can be tested, because the variable, the population and the reference value are named in it. H0 takes the absence of effect, H1 the expected direction.

Select the appropriate test statistic. The choice follows the shape of the data: comparing a mean with a reference value, comparing two proportions, comparing variances or testing how counts are spread across categories. The usual families carry the names t-test, z-test, F-test and chi-square test. The data team decides; the analyst states the conditions that choice assumes and checks that the data meets them. The guide notes that a properly formulated hypothesis does not require the underlying variable to follow a normal distribution.

Specify the level of significance. The sample size is fixed at the same time, from the smallest gap that would justify acting. A sample sized after the fact is a sample sized by the desired result.

State the decision rule. It names the critical value and the side of the distribution H1 points to.

Collect the sample and calculate. How the sample is drawn matters more than the arithmetic. The guide counts economy of collection among the technique's strengths: the formulation bounds the quantity of data needed to check a claim.

Make a decision, without amending the written rule. The result is stated in the protocol's terms: H0 rejected at the 5% level or H0 not rejected at that level.

Provide insights. Say what the result supports, what the measured gap is worth and what remains open.

What can be concluded from a result?

The p-value, as the American Statistical Association's statement defines it, is the probability of observing data at least as extreme as those obtained, assuming the null hypothesis and the statistical model chosen are true. It therefore does not give the probability that H0 is true, nor the probability that the data is due to chance: both readings invert the conditioning and are the most widespread misinterpretations.

A null hypothesis that is not rejected establishes that the data does not contradict it strongly enough at the chosen level. A sample that is too small produces the same verdict as a real equality. That is why a test that fails to reject H0 always comes with its power or with the confidence interval of the gap; otherwise the reader takes it as proof of equality. The guide makes this one of its limitations: results are probabilistic and are passed on with their significance level and their interval.

Statistical significance and practical significance

A statistically significant gap is one that chance in the sampling is unlikely to explain. Its size does not enter the verdict. The ASA statement makes a principle of it: statistical significance measures neither the magnitude of an effect nor its importance. With a large enough sample, a gap of CHF 2 per customer crosses the threshold as surely as a gap of CHF 200.

The business decision rests on two readings laid over each other: the first says the gap is real, the second compares the gap and its interval with the cost of the action it would trigger. A significant gap whose confidence interval covers the cost of the programme producing it does not justify a roll-out, and a non-significant gap on a pilot of thirty people does not justify dropping it.

The pitfalls

The threshold moved after the fact

The visible form is the threshold raised from 5% to 10% to save a result. The less visible form is more frequent: twelve segments, three channels and two periods are tested, and one of the seventeen tests comes out significant. At the 5% level and for seventeen independent tests, the probability that at least one crosses the threshold by pure chance is 58%. Two measures close this pitfall: the protocol declares which tests will be run, and the threshold is corrected for their number when the series is unavoidable, which is what the Bonferroni correction does.

A sample with no sampling rule

The test assumes a known sampling rule. A group of volunteers, a panel of survey respondents, the customers left once the others have churned: all three are put together in a way that produces the same statistical result as a real effect, when no effect is there. What the measured gap captures is then the way people entered the sample. The protocol names the target population and the sampling method; any divergence between the two is reported with the result.

Equality inferred from a null hypothesis that was not rejected

"No significant difference" arrives in the meeting as "the two options are equivalent", and the decision follows. The results note cuts that short by giving the interval: a gap between −CHF 20 and CHF 900 per case leaves the question open.

The test taken for the conclusion

The guide counts among its limitations that a test establishes the statistical soundness of a claim without standing in for analysis. Relying on it alone misses the signals a data exploration would have shown: a bimodal distribution, a dated break in a series. A test run without exploration answers a badly chosen question correctly.

A result nobody can read

The guide notes the difficulty of communicating the procedure and its result to stakeholders; it recommends building their confidence with simple examples. A one-page protocol, a reference value the business recognises and a single sentence stating the decision are worth more than a software output. Presenting the result then belongs to data storytelling.

AI considerations

The first use is formulation. Submitting a vague business claim and asking for three quantified H0 and H1 pairs, each with its variable and its population, forces what stayed implicit to be named. The model does the same job on the choice of test family: describing the shape of the data and asking which conditions the intended test assumes gives a list of checks that still have to be run.

The second use is execution: producing the R, Python or SQL code that computes the statistic and the interval, from a protocol already written. The third is translating the result for a committee. A model rewrites a software output into a single sentence stating the decision and proposes the simple example that makes the procedure understandable.

The model does not know how the sample was put together, which decides the validity of the test, and it will draw a conclusion from volunteer data without hesitating. It cannot set α, a trade-off between two business costs that only the organisation can quantify. Above all, an assistant asked which gaps cross the threshold in a data set will always find some: searching after the fact no longer costs anything. The only defence remains a protocol predating collection. Data sensitivity adds a further constraint: claims files or salary records carrying names must not go into a service hosted outside the organisation.

Examples

A health insurer in French-speaking Switzerland tests a diabetes prevention coaching programme on 220 members, before deciding whether to extend it to its whole portfolio. The protocol below is dated before the data extraction.

Protocol itemValue fixed before collection
Variable measuredAnnual claims cost, in CHF, per enrolled member
Null hypothesis H0The average annual cost of coached members equals the portfolio reference cost, CHF 4'800
Alternative hypothesis H1It is below CHF 4'800
Type of testOne-tailed, left tail, comparison of a mean against a reference
Significance levelα = 5%, critical value −1.65 (large sample, normal approximation)
Decision ruleReject H0 if the test statistic is below −1.65
Sample220 members enrolled in the programme, twelve months of claims
ResultValue
Average annual cost observedCHF 4'350
Gap from the reference costCHF 450
Standard deviation of the sampleCHF 900
Standard error of the mean900 ÷ √220 = CHF 61
Test statistic(4'350 − 4'800) ÷ 61 = −7.4
Decision−7.4 is below −1.65, H0 rejected at the 5% level
95% confidence interval of the gap, two-tailed by convention450 ± 1.96 × 61 = CHF 330 to CHF 570

The programme costs CHF 380 per member per year to roll out. The expected net saving is therefore CHF 70 per member, within an interval of −CHF 50 to CHF 190: the cost reduction is established at the chosen level, the net gain is not. A note that stops at the statistical verdict sends the committee off to roll the programme out across the whole portfolio.

The 220 members enrolled themselves. The protocol targeted the insured population and the collection supplied volunteers. A population that chooses a prevention programme looks after itself better to begin with. The defensible conclusion concerns the observed cost of those members. Attributing the gap to the coaching requires random assignment of participants. That divergence between the target population and the resulting sample is reported to the committee alongside the number.

Visualisations

Two parts of the technique call for a graphic form. The decision sequence is drawn, because its value lies in the order of the steps and in the barrier separating what is fixed before collection from what is calculated after. The pair of statistical significance and practical significance is drawn as a quadrant, because the two axes are independent and each of the four cells calls for a different action.

The rest is written as a table: the protocol carries one item per row, the matrix of the two errors sets the real state against the decision taken. A results note also gains from showing the confidence interval of the gap on a scale in francs, where the cost of the action sits as a marker; that representation belongs to business visualizations.

Cost

PhaseLevelJustification
PreparationMediumWriting the protocol takes a session with the business, where the reference value, the direction of the test and the threshold are negotiated.
ExecutionLowThe calculation is a spreadsheet function or one line of a statistics library. The cost sits in putting the sample together, which belongs to data preparation, upstream.
DocumentationLowThe protocol and the results note run to one page each. Archived once with the query that produced the sample, they need no upkeep.

Tooling

The protocol is written in a versioned document, and the tool matters less than the timestamp: one page in the project's document space, approved before the extraction, is enough to establish that it came first.

The spreadsheet covers the common cases. Excel and LibreOffice Calc offer t-test and normal-distribution functions, plus an analysis add-in for the z-test, the F-test and chi-square. They suit a single test on a sample already extracted and reach their limits as soon as the calculation has to be re-run at every data refresh.

Statistical environments are the data team's tool: R with its native functions, Python with SciPy and statsmodels. They give the statistic, the p-value and the interval in one command, keep a record of the calculation in a readable script and cover sample sizing. Power calculators such as G*Power do the same for anyone who does not program.

SQL builds the sample and the aggregates that feed the test, rarely the test itself. BI platforms display confidence intervals and uncertainty bands around a trend; they present a result obtained elsewhere, and an interval shown by default on a dashboard often reads as a test that was never formulated.

Sources

Growth-share matrix
All techniques
IDEF / IGOE