Paper Prototyping
Paper prototyping is a usability test. A representative user carries out a real task on a hand-drawn interface while someone from the team plays the computer: they swap the sheets in response to the user's finger, and they say nothing. BABOK files it among the prototyping methods and describes it in a single line, paper and pencil to draft an interface or a process, which leaves out the thing that makes the technique: the session. What the session returns is observed behaviour, the point where the finger hesitated, the path the interface did not have, the word nobody understood. The correction is made in pen, on the sheet, between two participants, and that speed of correction is what you are buying.
Goal
Paper prototyping establishes, in an hour and in the form of observed behaviour, whether a chosen but not-yet-built design holds up on contact with a real user: what the user did, where the finger stopped, which word they misread, which path they looked for and the interface did not have. The two other routes to that answer cost more or return less: asking a stakeholder gives only an opinion, building it answers only at the price of a sprint.
The question it settles has a precise shape: does this design survive contact with a real user, and if not, exactly where does it give way? The answer arrives while a change still costs a pen stroke.
The deliverable is two objects. The log of observed problems: where the participant hesitated, the paths the design did not have, the words they misread, the steps they invented. And the corrected deck: the sheets reworked in pen between two sessions, together with the design decisions those pen strokes record. The paper itself is thrown away.
The economics of the technique is its entire argument, and Rettig stated it in 1994: low-fidelity prototyping works because it maximises the number of times you get to refine your design before you must commit to code. The benefit rests on this: a cheap artefact can be corrected between two participants, so one afternoon carries three rounds of design where a clickable prototype carries one. The Nielsen Norman Group reaches the same place from the other end: for the same budget, three studies of five users beat one study of fifteen, because what you are buying is the redesign done between rounds.
BABOK separates the approach from the method, and the two are free of each other. The approach says what becomes of the prototype: it is thrown away (throw-away prototyping), or it is grown into the delivered solution (evolutionary prototyping). The method says what it is made of and what someone does with it: you read a storyboard, you operate a paper prototype, you execute a simulation. Every prototype carries an answer on each of the two axes. Paper prototyping is a method, and on the approach axis it occupies a single cell: paper does not ship, so a paper prototype is always a throw-away prototype. Routing between the two axes belongs to prototyping taken as a whole.
Usage
When to use it
- Design settled, nothing coded: the only question left is whether a real user can operate it.
- Interface wording in dispute: labels, error messages and terminology are exactly what the session tests.
- Team disagreement on navigation: the participant's finger settles it, and nobody has to be right.
- Correction expected between two participants: the sheet is reworked in pen, so three rounds fit into one afternoon.
- Non-technical stakeholders in the room: the observer bench lets them watch the failure, which no written report achieves.
- No tooling available: the technique presupposes no licence, no design system, no environment.
- Team gathered around one table: the deck is a physical object, and the session lives on it.
When not to use it
- Fine-grained or continuous interaction (dragging, scrolling, animation): the human computer cannot keep up, reach for an interactive high-fidelity prototype.
- Design still open: there is nothing to operate, reach for a throw-away mock-up or a storyboard.
- A question of volume, delay or cost: the session returns no numbers, reach for simulation.
Description
Preparing the deck
The tasks are written before the first sheet
The task list decides which sheets have to be drawn, and the reverse order produces a demo: you draw the screens you like, then invent the tasks that suit them, and the session confirms what the team already believed. It is the commonest preparation defect, and it is invisible from the inside.
The background, then the movable pieces
The fixed chrome is drawn once, on one sheet: the window, the header, the persistent navigation. Everything that changes in response to the user becomes a separate, physically detachable piece: index cards for dialogs, menus and dropdown lists, sticky notes for tooltips and error messages, paper slips for the content of a field. This separation is what makes the interface operable: without it, what is left is a drawing.
An acetate sheet and a wipeable pen
Laid over the sheet, they let the participant type into a field, and the sheet survives intact into the next session.
Dummy text where the content is not under test, real text where the word is the test
Labels, button text, error messages and business terminology are precisely what a paper prototype finds problems in, so they are written for real. A label left as lorem ipsum removes from the session the very thing the session sees best.
Hand-drawn
BABOK gives the reason: faced with a throw-away or paper mock-up, users may feel more comfortable being critical of it because it is not polished and release-ready. A printed, pixel-aligned screen invites politeness. Roughness invites criticism, and criticism is the product.
The deck is ordered
It is ordered so that the human computer can find any piece in one second: by screen, by state, with the error messages kept apart. The Nielsen Norman Group names disorganisation as a failure mode of the technique, and the mechanism is this: the computer rummages, the rummaging inserts a pause, the pause breaks the participant's train of thought and the data degrades.
Two pen colours
One for the drawing, one for the corrections made during and between sessions. The correction history is a result in its own right, and it stays visible on the sheet.
The task script
Tasks, not questions
A task states a goal and a starting situation, and it never names the interface element that achieves it. "Book the same slot as last week" is a task. "Tap My appointments" is an instruction, and a task that names the control has already answered its own question.
The tasks are drawn from real work, the work described by the business use cases and scenarios or by the user stories already written. They are ordered from simple to complex, and three to six tasks fill an hour. One card per task, handed over one at a time, so that the participant meets each situation at the moment they have to deal with it.
The four roles
The participant, the facilitator, the human computer, the observers. The technique is played out here, on what each of them forbids themselves.
The participant is a representative user: someone whose job is the one the interface serves, recruited from outside the team and the project. They carry out the task and they think aloud, continuously.
The facilitator hands over the task cards one at a time, keeps the verbalising going and runs the debrief. Rettig states the rule: never tell the user how to do it. To the question "what do I do now?", the facilitator returns the question: "what would you do?".
The human computer is a single person. They manipulate the paper: they swap the sheet, lay down the dialog, reveal the error message, in response to the finger and to nothing else. They say nothing. Snyder rests the whole definition of the technique on that exact point: the person playing computer does not explain how the interface is intended to work. If the participant touches a place for which the design has no answer, the human computer does nothing, and that silence is the result: a missing path has just been measured, and no opinion would have flagged it.
The observers, meaning the rest of the team (designers, developers, product owner), log. Rettig gives a mechanical discipline: one problem per card. They do not speak, do not defend the design and do not answer a question addressed to the room.
Who sits on the observer bench is decided before the session
The technique asks one person to fail out loud in front of the team that designed the thing, and it asks that team to watch without rescuing. The bench is therefore the project team. The participant's own management is kept off it: an employee operating the interface under their line manager's gaze stops exploring and starts succeeding, and the data is dead before the first task card is handed over. When the sponsor wants to watch, they watch the recording or follow the session from an adjoining room. In an organisation where contradicting a senior manager out loud is not done, it is the composition of the room, more than the script, that decides what the session will return.
A fifth hat is worn by any one of them, without adding anyone to the room: the greeter, who receives the participant, takes their consent and runs the preamble.
Running the session
- The preamble
"We are testing the design, not you. No action is wrong. If something is confusing, the design is at fault." Then consent and confidentiality. The paper and the human computer are explained in two sentences, once and without apology. - A warm-up task
A trivial one, so that the participant gets into the habit of speaking while doing. Nobody thinks aloud spontaneously. - Thinking aloud
Nielsen describes it this way: you ask the participant to use the system while continuously thinking out loud, that is, simply verbalising their thoughts as they move through the user interface. The facilitator prompts only to keep the verbalising going ("what are you thinking?"). Nielsen's warning is the one to carry: interruptions and clarifying questions very easily change user behaviour. A badly run session then measures the conversation. - The tasks, one at a time, in a silent room
The participant reads the card, puts it down and works. - Log the behaviour
What the participant did, where the finger hovered before landing, what they said while doing it, what they expected and what happened instead. The question "do you like this screen?" returns an opinion, and opinions are collected at the debrief at the end of the session, where they weigh what an opinion weighs.
The loop between two participants
The session returns the behaviour, the loop converts it into a corrected design and both are the technique. Immediately afterwards, the team sorts the observation cards, groups what recurs and agrees the changes together, while everyone still has the scene in mind. The change is made on the sheet in pen, or the sheet is redrawn in ten minutes, and the next participant tests the corrected deck. No clickable prototype turns at that speed.
Snyder states the limit of the exercise: watching only one test is better than not watching any at all, but if the test turns out to be atypical, it can skew the observer's view of what needs to be changed. Three to five participants, with a correction between each, is the usual shape. Nielsen puts numbers on the rest: five users per study and three studies of five beat one study of fifteen.
Scope and limits of the session
Snyder lists what comes back: most usability problems, including issues with concepts, terminology, navigation, workflow and screen layout. Add to that the step the user expected and the design does not have.
Out of reach are aesthetics, colour and graphic treatment, which do not read on paper as they read on a screen; perceived performance and response times, which depend on a machine; and any rapid, subtle or continuous interaction, which the human computer cannot reproduce. The interactive high-fidelity prototype takes over on those questions, and saying so honestly is what keeps the technique from over-promising.
The pitfalls
The room rescues
The human computer lets slip "no, actually you would tap there", or the facilitator points. Within the second, the session has stopped measuring the interface and started measuring the explanation. Silence is the instrument, and it is held by discipline.
Questions instead of tasks
A question returns an opinion, a task returns behaviour, and behaviour is the technique's only product.
Testing with colleagues
They know the domain, the vocabulary and the intended flow, so they do not stumble where a real user stumbles. The session then returns the team's own mental model, laundered by a third party, and the team is reassured by it.
The corrected deck goes out as the specification
BABOK describes the mechanism: stakeholders focus on the design specifications rather than on the requirements any solution has to address, and developers believe they must reproduce the prototype exactly. On a paper prototype that takes a precise shape: the corrected sheets go to the developers as the spec, and the requirement never gets written. What the session established is a requirement, "the user must be able to find an existing appointment from the home screen"; the sticky note on the sheet is one possible realisation of it. The first is written down and tracked. The second is thrown away with the paper.
The deck with no task script
"It is only paper", so nothing to write in advance. A deck without tasks is a demo, and a demo is a presentation in which the designer holds the pen.
A disordered deck
The computer rummages, the pause breaks the participant's train of thought, and the hesitation that gets logged belongs to the paper.
A prototype that is too neat
Printed screens invite politeness, and they also invite the debate about the typeface, which is not the question of the day.
Rewriting the design after one participant
A sample of one, in which the atypical weighs as much as the recurring.
Nobody logs
With no cards, the loop has no input, and the session was well-acted theatre.
AI considerations
Two limits specific to this technique govern all the rest.
A language model cannot play the computer. The entire function of that role is to withhold: respond to the finger, explain nothing, repair nothing. A model asked to hold the interface will explain, help, complete and forgive, which is to say it will produce exactly the defect the role exists to prevent. The instrument is a person who says nothing.
A language model cannot be the participant. A synthetic user does not stumble where a real one stumbles, because it has, in effect, already read the design. A "simulated user test" returns the model's prior, and the only product of this technique is observed behaviour. A paper prototype tested against a language model has produced nothing at all.
Where tooling earns its place is around the session. Upstream: rough out the sheets from a requirement or a user story, leaving the team to redraw them by hand afterwards, roughness being a wanted property; audit the task script by asking explicitly whether any task names an interface control, which catches the commonest writing defect; fill the sheets with plausible Swiss content (names, addresses, four-digit postcodes, dates, times, amounts in CHF) so that a field reads like a real field. Downstream: transcribe the observation cards, cluster the recurring stumbles, draft the change list. It is thankless, it is where the hours go, and it delegates without harm.
The data. A session filmed in a practice, with a real patient, produces sensitive personal data under the revised Federal Act on Data Protection (revFADP), because it is health data. Explicit consent and no session recording lodged with a provider without a data-processing agreement.
Examples
A physiotherapy practice is putting appointment booking online. The screen is drawn, nothing is coded. A patient of the practice comes in for an hour: one colleague plays the computer, a second facilitates, a third logs. Five tasks, five cards, handed over one at a time.
| Task handed over | What the participant does | What the prototype does | Correction in pen |
|---|---|---|---|
| Book the same slot as last week | Touches the physiotherapist's name on the home sheet, expecting her diary | Nothing: no sheet matches | Sticky note "My appointments" added to the home sheet |
| Book a 30-minute session | Reads "Short session", stops: "how short is short?" | The next sheet is laid down, the hesitation is logged | Label replaced by "30-minute session" |
| State that the session is prescribed by the doctor | Looks for a field for the prescription, scans the sheet twice | Nothing: the field does not exist | Field "Doctor's prescription" added below the reason for consultation |
| Cancel Thursday's appointment | Opens the menu and looks in it for "Cancel" | The menu is laid down, it does not contain "Cancel" | "Cancel" moved onto the appointment's own card |
| Confirm the booking | Touches "Confirm", then waits, hand in the air: "do I get something?" | The confirmation sheet is laid down | Note "Confirmation sent by email" added to that sheet |
The five rows say what the participant did. That is the material the technique produces; what she thinks of the screen is collected at the debrief, separately, and weighs what an impression weighs.
The two rows where the third column says "nothing" are the most valuable of the session, and they exist only because the human computer stayed silent. Someone who had answered "it is in the menu, at the top" would have erased the discovery on the spot, and the log would have returned three rows instead of five.
The five corrections are made in pen, on the existing sheets: one sticky note, two labels, one field added, one control moved. The next participant tests the corrected deck straight away, and so tests a design the previous session has already changed.
Visualisations
Three figures carry this technique, and each answers a different question. The crossing of the two axes places the technique among its neighbours: a method and a single cell on the approach axis. The session shows the room and the loop that runs through it, from the participant's finger to the swapped sheet, from the logged hesitation to the sheet reworked in pen and on to the next participant. The session log shows what all of that produces: five tasks, five observed behaviours and the corrections they triggered.
The session figure reads back as a checklist, and a real session has to answer yes to its four questions. Does the participant come from outside the team? Is the human computer one person, and does that person stay silent? Is someone logging, one problem per card? Was the sheet reworked before the next participant? A session missing that last loop has produced a list of problems, where the technique promised a corrected design.
Cost
| Phase | Level | Justification |
|---|---|---|
| Preparation | Medium | The design has to be settled, the tasks written, the deck drawn and then ordered so that the human computer finds any piece in one second. The heavy item is recruiting representative participants, from outside the team and available for an hour each. The materials cost a few francs and are prepared in one afternoon. |
| Execution | Low | One hour per participant, three to five participants, three or four people from the team around the table. No environment to prepare, no licence, no deployment, and the cards are sorted within the half hour that follows the session. |
| Documentation | Low | The product fits into a log of observed problems and a corrected deck, photographed state by state at the end of the session. There is no model to maintain over time: the corrections that become change requests go to item tracking, and the paper is thrown away. |
Tooling
The kit fits in a box: A3 or A4 sheets for the backgrounds, index cards for dialogs, menus and lists, sticky notes for tooltips and error messages, paper slips for the content of a field, scissors, blank labels or correction fluid, two pen colours. An acetate sheet laid over the paper, with a wipeable pen, lets the participant type into a field without damaging the deck. A camera closes the session: each state of the deck is photographed, because the deck is the archive.
For a distributed observer bench, a document camera (or a phone on a gooseneck mount) and a video call are enough to show the table. The session loses by it: the human computer's latency rises, the participant's finger becomes a cursor, and the observers stop reacting together. That is weighed against the cost of travel.
The session log lives in a table or a spreadsheet, one row per observed problem. The corrections that become change requests then go to item tracking, which exists for that.
Interactive mock-up tools (Figma, Axure and their like) are the exit from the technique. As soon as the sheets become clickable, the human computer disappears, and with it the silence that was the instrument. You move there when the interaction becomes too fine-grained for paper, and that move is a change of technique, decided as such.
Sources
- IIBA, A Guide to the Business Analysis Body of Knowledge (BABOK Guide) v3, §10.36 Prototyping: paper prototyping as one of the prototyping methods, the separation of the approach axis from the method axis, the observation that users feel freer to criticise a mock-up that is not polished and release-ready and the risk of the design taking the place of the requirement.
- Carolyn Snyder, Paper Prototyping: The Fast and Easy Way to Design and Refine User Interfaces, Morgan Kaufmann: the definitive treatment, the definition of the technique as a usability test in which a person plays the computer without explaining the interface, what the session finds, what it leaves and the limit of the single session.
- Marc Rettig, Prototyping for Tiny Fingers, Communications of the ACM 37(4), April 1994, pp. 21-27: the origin of paper prototyping, the roles of the session, the facilitator's rule never to tell the user how to do it, one problem per card and the economic argument for the technique.
- Jakob Nielsen, Thinking Aloud: The #1 Usability Tool, Nielsen Norman Group: the think-aloud protocol and the warning about interruptions that change user behaviour.
- Nielsen Norman Group, UX Prototypes: Low Fidelity vs. High Fidelity: the limits of the human computer, the part played by an ordered deck and the point at which the interactive high-fidelity prototype takes over.
- Jakob Nielsen, Why You Only Need to Test with 5 Users, Nielsen Norman Group: the number of participants per study and the fact that three studies of five beat one study of fifteen.

