Simulation
A simulation is a prototype you execute. You give it a process, scenarios, business rules, data and inputs, and it produces a behaviour whose measures you read off: throughput, queue length, cycle time, utilisation, cost, delay. BABOK places it among the prototyping methods and writes that it is used to demonstrate a solution or components of a solution, and that it may test processes, scenarios, business rules, data and inputs. What decides its value is its behavioural fidelity, demonstrated in one way only: by reproducing a period whose outcome the organisation already knows. A simulation nobody has validated remains a plausible fiction, a fiction with decimal places.
Goal
Simulation answers a question of a precise shape: if we proceed this way, what happens and by how much? It answers before anything is built or changed, in numbers, in the very unit in which the decision is taken. The other prototyping methods produce something to look at, a sequence of frames or an interface drawn by hand, which the stakeholder comments on. Simulation produces something to execute, whose results the stakeholder reads.
BABOK describes the technique in two sentences, each carrying a reservation that has to be kept as it stands: simulation is used to demonstrate a solution or components of a solution, and it may test processes, scenarios, business rules, data and inputs. The vocation of measurement is anchored elsewhere, in one place only: in the task Measure Solution Performance, where BABOK writes that prototyping is used to simulate a new solution so that performance measures can be determined and collected. That is the only place in the guide where prototyping is tied to measurement. The guide writes it of prototyping in general, through the verb "simulate". The only method in the family that accomplishes it is this one. The two anchors say different things worth holding apart: one says what the technique exercises, the other what it returns.
The deliverable consists of seven objects. The executable model. The parameter sheet, which carries the provenance of each one: measured, fitted or assumed. The validation record, which confronts the model with a period whose actual results are known and states the tolerance adopted. The scenario definitions. The run results, with their intervals. The sensitivity analysis, which names the parameter the decision actually rests on. And the decision note, which states the result's domain of validity, that is, the conditions under which it was established and outside which it ceases to hold.
What the organisation buys at that price is the possibility of being wrong for free. A dunning policy, a tariff band, one more counter, a routing rule: applied for real, these changes cost their implementation, their reversal and the harm done in the meantime. Executed in a model, they cost a night of computation. BABOK observes, in connection with defining the future state, that prototyping could also help determine the potential value of an option; simulation is the method that determines it by measuring it.
Usage
When to use it
- A rule change with countable consequences: dunning, tariff bands, eligibility criteria, routing rules.
- A capacity or throughput question: queues, waiting times, utilisation, one counter more or one counter less.
- Options that cannot be tried in parallel: two dunning policies on the same customers, impossible in reality.
- A change that is expensive or hard to undo: the run is the cheapest way to be wrong.
- Performance measures expected before go-live: the promised value is quantified while it is still possible to walk away.
- A process model already exists, the commonest trigger: making it executable costs little.
- Counter-intuitive system behaviour: a bottleneck that moves, a queue that explodes past a threshold.
When not to use it
- The decision costs less than the model: make the change and measure reality, through a pilot, a limited rollout or an A/B test.
- The question is about the appearance of the solution: there is nothing to execute, so take paper prototyping or storyboarding.
Description
Two independent axes
BABOK classifies prototypes on two distinct elements. The approach says what becomes of the prototype: you throw it away, or you grow it into the solution. The method says how it is made and exercised: through a storyboard, on paper, through a simulation or through workflow modelling, which appears on that axis as a borrowed method whose home is process modelling. Every prototype answers on each of the two axes, independently. A paper prototype is always a throw-away prototype. A simulation may be thrown away or live on in a tool that keeps it alive, which places it on the side of evolutionary prototyping. Asking "is it throw-away or is it a simulation?" is asking a question with no answer: it is prototyping taken as a whole that governs the routing between the two axes.
The verb simulate runs through this family and creates a particular confusion in it. Several prototype forms can simulate a process or a rule, and the guide says as much of the evolutionary tool, provided that specialised software is used: that is a capability of the tool. Simulation is the method for which that is the entire purpose and the only product. The question that isolates it from the others fits on one line: does the prototype have to be executed or only looked at? Executed, it is this one. Looked at, it is a storyboard or a paper prototype.
Running a simulation
Frame the question before a model exists
Name the decision, then the measures that settle it. This is conceptual modelling: choosing what is modelled and what is left out. Robinson calls it the most difficult, least understood and probably most important activity in a simulation study. A model built with no decision attached produces a number nobody uses. Detail is justified by the decision: every element added must move a measure that matters, failing which it adds a parameter to estimate and a source of error to carry.
Start from the existing process model
A simulation executes on a process model, produced by process modelling, with its own notation and its own rules. Simulation loads that drawing with what makes it executable: durations on the activities, probabilities on the branches, resources with their capacities and their schedules, queue disciplines (first come first served, priority, abandonment after so many minutes) and an arrival law for the cases. A drawn model is a map; the same boxes, fitted with those five families of parameters, become a machine.
Make the rules executable
A rule stops being a sentence you read and becomes a condition the run evaluates. "The customer is normally chased within the month" does not execute: you need the threshold in days, the order in which the conditions are evaluated, the fate of the boundary case and the exception everyone applies without ever having written it down. This is the work of business rules analysis. Simulation exposes the rules left vague: an ambiguous rule cannot be coded, and the modeller discovers the ambiguity at the moment of writing it.
Parameterise on real data and record the provenance of every parameter
Every parameter in the model is measured (it comes out of a system and can be shown), fitted (a distribution has been calibrated on a history) or assumed (nobody has observed it, somebody judged it). An assumed parameter is a hypothesis and carries that name on the parameter sheet. Sargent calls this requirement data validity: ensuring that the data needed for building the model, evaluating it and running the experiments are adequate and correct. NASA-STD-7009B makes it an object of assessment in its own right, input pedigree, the credibility of a result depending first of all on the credibility of what went into the model. The failure mode is banal: an assumed parameter that looks like a measured one, because nothing in the table tells them apart.
Verify, then validate
These are two different questions, and the first is the easier. Verification asks whether the program does what was specified: it establishes that the computerised model and its implementation are correct. Validation asks whether the model is right for the use being made of it: Sargent defines it as the substantiation that a model, within its domain of applicability, possesses a satisfactory range of accuracy consistent with the intended application. A model without a single bug can be wrong from end to end, because it faithfully represents a process that is not the organisation's. The instrument of validation is the reference scenario: the model is executed with the rules already in force, over a period for which the organisation holds the actual results, and the model is accepted only if it reproduces them within a tolerance fixed before the run. Fixed before, because a tolerance chosen after the run is chosen so that it will be met.
Replicate
A stochastic model returns a different answer on every execution, and the gap between two executions is anything but marginal. Law gives the cheapest demonstration of it: on a bank-teller model, five independent replications of the same scenario produce average queue waits of 1.53 · 1.66 · 1.24 · 2.34 · 2.86 minutes, and he concludes that one run clearly does not produce "the answers". The mode of operation he describes as common, a single run of arbitrary length whose result is treated as the true characteristic of the model, could expose the analyst to a significant probability of making erroneous inferences about the system under study. The number of replications is therefore decided in advance, by the precision wanted on the measure that settles the decision. Law gives a mechanical procedure: run a pilot batch, ten replications or so, compute the half-width of the confidence interval on the deciding measure, and add replications until it falls below the precision the decision needs. A decision that turns on five cut-offs demands an interval narrower than five cut-offs. That calculation fixes the number, and the twenty replications of the example come out of it.
Know whether the model terminates or runs in steady state
A simulation with a natural beginning and end, a billing year for instance, is called terminating: the system and the model start empty, and each replication replays the whole period. A steady-state model is asked about the behaviour of the system once it has settled, and it too starts empty, queues empty and counters idle, in a state foreign to the one that is to be measured. The warm-up period is the time it takes to leave that transient, and the observations it produces are deleted before anything is computed. The classic error is to read an average waiting time into which an hour of empty queues has been folded: the measure turns optimistic, the more so the shorter the run. Capacity and queue questions, the commonest trigger for the technique, almost always fall on the steady-state side. The dunning model in the example is terminating, and the warm-up question does not arise for it.
Read the result with its uncertainty and find the parameter that carries the decision
A simulation result is an interval: it is reported with its bounds. Then comes the sensitivity analysis: the assumed parameters are varied, one at a time, across the range in which they remain credible, and you observe which of them move the deciding measure. Often only one does, and the exercise has rendered its best service: it has said what the decision rests on and where it would be worth going to measure rather than continuing to assume. NASA-STD-7009B treats this whole apparatus as an engineering requirement and separates two credibilities, that of the model and that of the results it produced, requiring for the latter the characterisation of uncertainty, sensitivity analysis and a proper reporting of results.
Decide and record the domain of validity
Sargent puts it plainly: a model may be valid for one set of experimental conditions and invalid for another. The domain of validity is therefore part of the deliverable, on the same footing as the number. It states what the model was validated to do, on what data and within what bounds of volume and behaviour. It will prevent, six months from now, a model validated for a capacity question from being asked a pricing question.
What makes a simulation fail
The plausible model nobody validated
It runs, it produces numbers, the numbers have decimal places, and nobody ever checked that it reproduces a year everyone knows. This is the central trap of the technique, and the guard against it is structural: the reference scenario and a tolerance settled before the run.
The false precision of a single execution
One run of a stochastic model is a draw: taking it for the answer is the error.
Confusing verification with validation
The model does exactly what it was asked to do, and what it was asked to do is wrong. A review of the model reassures on the first question and is silent on the second.
An assumed parameter that carries the whole decision
It has never been observed, it is written like the others, and the result swings with it. Sensitivity analysis is what flushes it out.
The model reused outside its domain
Validated for a capacity question, asked six months later about pricing. It answers with no guarantee.
The theatre of fidelity
Polishing the detail of the model while its inputs remain assumed. BABOK notes that, if the system or process is highly complex, the prototyping process may become bogged down in discussion of "how" rather than "what": the debate about the model's granularity takes the place of the debate about the quality of the data.
AI considerations
One remark, belonging to this technique, governs all the others. Asked what would happen if the dunning rules were tightened, a language model will return a result: fluent, plausible, confidently phrased, along the lines of "cut-offs would rise by about 5%". It executed nothing. It produced a supposition, in the register of a result. Simulation exists to replace that sentence with a measured one. A language model narrates an outcome, a simulation computes it, and confusing the two is all the easier because the narrated version reads better and arrives in three seconds.
Where tooling renders a real service, it renders it upstream and downstream of the run. Upstream, fitting the input distributions to a billing history or an event log is a tedious piece of statistical work that the machine does well, and it does something better still that the practitioner neglects: flagging the series too thin for anything to be fitted to them, which is exactly the input-pedigree problem. Building the skeleton of the model from an existing process model or an event log, through process mining, saves days. Generating the scenario sweep avoids writing twenty variants by hand. Downstream, a surrogate model learned on the simulation's own runs sifts a large scenario space cheaply, reserving the expensive executions for the survivors, and drafting the sensitivity analysis can be delegated without harm.
Three things remain beyond the tool's reach. Behavioural response is a hypothesis, and an AI will supply one with ease: the way customers react to a rule that has never been applied is what no historical data contains, and an invented elasticity has the same tone as a measured one. Validation requires a referent, that is, a known reality against which to confront the model: no tool possesses it, and it is the organisation that holds it in its own systems. Data sensitivity, finally: a billing history identifies real people, and it is pseudonymised before it enters a model, all the more so before it leaves for a third party.
Examples
A municipal utility is revising its dunning rules. The question the run has to settle: do the proposed rules collect more and at what price for the residents?
| Parameter | Value | Provenance |
|---|---|---|
| Volume of the reference year | 12'000 invoices, CHF 4'800'000 billed | Measured (billing system) |
| Rules A, in force | 1st reminder d+30 · 2nd reminder d+60 (CHF 20 fee) · cut-off notice d+90 · cut-off d+105 | Measured (regulation in force) |
| Rules B, proposed | 1st reminder d+20 · 2nd reminder d+35 (CHF 20 fee) · cut-off notice d+50 · cut-off d+65 | Defined (draft regulation) |
| Payment delay | distribution fitted on twelve months of receipts | Fitted |
| Customer response to the tightening | payment moves forward in the wake of the reminder, the default rate is unchanged | Assumed |
| Replications per scenario | 20 | Run choice |
| Acceptance tolerance | 5% on each measure | Fixed before the run |
| Measure | Observed (past year) | Scenario A: rules in force, simulated | Gap A / observed | Scenario B: rules proposed, simulated |
|---|---|---|---|---|
| 1st reminders issued | 2'160 | 2'147 (± 31) | 0.6% | 3'480 (± 44) |
| 2nd reminders issued (CHF 20 fee) | 604 | 611 (± 18) | 1.2% | 742 (± 22) |
| Cut-off notices | 138 | 141 (± 9) | 2.2% | 165 (± 10) |
| Cut-offs executed | 41 | 39 (± 5) | 4.9% | 44 (± 5) |
| Average payment delay | 38.4 days | 38.1 days (± 0.4) | 0.8% | 31.6 days (± 0.5) |
| Collected at 90 days | CHF 4'603'000 | CHF 4'598'000 (± 14'000) | 0.1% | CHF 4'690'000 (± 16'000) |
| Reminder fees billed | CHF 12'080 | CHF 12'220 (± 360) | 1.2% | CHF 14'840 (± 440) |
The model is weakest where the counts are smallest: the cut-offs clear the tolerance at 4.9%, and only just, so a decision that turns on cut-offs rests on the thinnest cell in the model.
The proposed rules collect CHF 92'000 more at 90 days and shorten the average delay by 6.5 days, at the price of 1'333 first reminders and 131 second reminders more and of 5 further cut-offs in residents' homes. Pricing that trade-off in francs would require a unit cost per reminder and a social cost per cut-off, which this run did not produce.
Visualisations
The technique produces two objects, each taking the form that follows from its nature. The run is a spatial object: inputs converging, a model executed a certain number of times, measures coming out of it with their intervals and above all a return edge, the one that runs back from an observed reference to the model and governs everything else. That edge, belonging to simulation alone, is drawn: a three-row table would destroy it. The parameter sheet and the run table are rows and columns, with a provenance column and a gap column that read at a glance: they stay in cells.
Three checks are enough to audit a real simulation study. Does every input carry a declared provenance, or did some arrive without anyone knowing where from? Does the replication counter show a number greater than one? And has the return edge been traversed, in other words has the model been confronted with a period whose outcome was already known? A study in which that edge is missing stopped before the step that made its numbers readable.
Cost
| Phase | Level | Justification |
|---|---|---|
| Preparation | High | Preparation is where the technique's cost lives, and that is what separates it from the other prototyping methods. Framing the question, the process model, writing the rules executably, extracting and fitting the data, then validating against a known period: validation belongs to preparation and is the heaviest item in it. Reckon on weeks, more if the data have to be measured before they can be fitted. |
| Execution | Low | Execution consumes processor time. Twenty replications of two scenarios run overnight, often in a few minutes, and sweeping a scenario space is paid for in machine time rather than in workshops. This is the technique's economic inversion: expensive to build, almost free to interrogate, which makes it pay as soon as there is more than one question to put to it. |
| Documentation | Medium | The parameter sheet with the provenance of each entry, the validation record with its tolerance, the scenario definitions, the results with their intervals, the sensitivity analysis and the domain of validity. None of that can be improvised after the fact, and the domain of validity is what decides whether the model lives six months or six years. |
Tooling
The simplest tool is the spreadsheet, and it is worth knowing what it can do and where it stops. It carries a deterministic model very well: volumes, rates, average durations, and the effect of a rule change is read off it. That is where the majority of the things called simulations live, and it is legitimate as long as the question does not depend on variability. As soon as there is a queue, a shared resource, an abandonment or a heavy-tailed distribution, the average becomes misleading, because the behaviour of a resource-constrained system does not follow from the average of its inputs. The spreadsheet then stops answering the question that was asked.
Above it come the process simulation engines, which take a process model that has already been drawn and add the parameterisation that makes it executable. The WfMC's BPSim specification exists for exactly that: it defines an interchange format that attaches scenarios and their parameters (durations, resources, costs, probabilities, arrival laws) to a BPMN or XPDL model, which allows parameterisation in one tool and execution in another. Several business-modelling and process-mining platforms implement it, which makes it the cheapest path when the process model already exists.
For capacity and queue questions, the discrete-event simulation tools (Simul8, Arena, AnyLogic, SimPy for those who program) are the right level: they natively carry resources, queues, priorities, schedules and probability laws, and they return confidence intervals. They are the ones that know how to conduct the replications and the output analysis without being rewritten. The process-mining platforms play a complementary and underrated role: they reconstruct the real process from the system's event logs, which yields at once the model and part of the measured parameters, and they substitute the executed process for the declared one.
The run can also do without a machine. The ABPMP's BPM CBOK puts the alternative in two terms: a simulation is either manual or, when it runs through a process simulation tool, electronic. It describes the first term under the name process laboratory, where a small cross-functional team walks mock transactions by hand from one end of the process to the other, inside a process improvement, redesign or reengineering effort.
The choice of tool finally weighs far less than the quality of what is poured into it. A simulation run in a spreadsheet on measured data and validated against a known year is worth more than a simulation run in the best engine on the market on assumed parameters. The tool decides what the model can represent; the data decide what its results are worth.
Sources
- IIBA, A Guide to the Business Analysis Body of Knowledge (BABOK Guide) v3, §10.36 Prototyping: simulation as a prototyping method, what it demonstrates and what it may test, as well as the limitation about becoming bogged down in the "how".
- ABPMP, Guide to the Business Process Management Common Body of Knowledge (BPM CBOK), §3.10.2 Simulation Tools and Environments: manual or electronic simulation and the process laboratory, where a small cross-functional team executes mock transactions by hand from end to end.
- Robert G. Sargent, Verification and Validation of Simulation Models, Proceedings of the 2010 Winter Simulation Conference, pp. 166-183: verification, validation and data validity as three distinct questions and validity always relative to the intended use.
- Averill M. Law, Statistical Analysis of Simulation Output Data: The Practical State of the Art, Proceedings of the 2015 Winter Simulation Conference, pp. 1810-1824: replications, run length and confidence intervals and the quantified demonstration that a single run produces no answer.
- NASA-STD-7009B, Standard for Models and Simulations, 2024: the credibility of a model and that of its results treated as two separate assessments, with input pedigree, uncertainty characterisation, sensitivity analysis and the reporting of results.
- Stewart Robinson, A Tutorial on Simulation Conceptual Modeling, Proceedings of the 2017 Winter Simulation Conference, pp. 565-579: framing the model, what is modelled and what is left out.
- WfMC, Business Process Simulation Specification (BPSim) v2.0: the interchange format that parameterises a BPMN or XPDL process model to make it executable.

