Card Sorting
Card sorting is a user-research method. Participants are given a deck of cards, each carrying an existing piece of content, a page, a service, a form, then they group them the way that makes sense to them and name the groups. Those groups and their labels expose the mental model by which these people expect to find the material. The result feeds an information architecture: a navigation menu, a taxonomy, the outline of a document. A sort produces a candidate structure, which is then put through a separate check.
Goal
In a card sort, users group existing pieces of content into the categories that feel natural to them, then they name those categories. The method answers a design question: how does the audience expect to find this material, when the vocabulary and the divisions of the organisation that publishes it are not imposed on them?
It supports the choice of an information architecture: the top-level sections of a site, the order of a menu, the entries of a document taxonomy. The deliverable has two halves: on one side the groups as the participants formed and labelled them, qualitative material about vocabulary; on the other the count of card pairs filed together, quantitative material about the strength of each association. The two are read together: when a firm association attracts diverging labels, the group holds and the work is on the label.
Measuring whether people find what they are looking for in the resulting structure calls for tree testing.
Card sorting and the affinity map
The two techniques look alike from a distance, cards moved by hand until groups appear; the direction of the exercise separates them.
| Who sorts | What is sorted | What comes out | |
|---|---|---|---|
| Card sorting | users or stakeholders from outside the team | pieces of content that already exist | the organisation those people see in it, the basis of an information architecture |
| Affinity map | the team itself | the ideas it has just produced in session | convergence on a few themes after a phase of divergence |
If the material to be sorted was written within the hour by the people in the room, the affinity map is the tool; if it already exists and the question is how an audience files it, this is a card sort, with participants recruited from outside the team.
Card sorting settles the grouping. Settling the meaning and the vocabulary of the terms belongs to concept modelling and to the glossary; eliciting one by one the criteria by which a person tells two items apart belongs to the repertory grid.
Usage
When to use it
- Redesign of an existing navigation: the content inventory is available and the current tree is contested.
- Internal vocabulary far from the public's: the site is divided along the org chart and visitors do not recognise the labels.
- A candidate taxonomy to put to the test: a closed sort, whose categories are fixed in advance, shows which of them take content without hesitation.
- Internal disagreement about classification: the sort moves the arbitration from departmental opinion to user data.
- A dispersed audience: an unmoderated remote sort recruits widely at a low marginal cost per participant.
When not to use it
- A structure to evaluate rather than to produce: use tree testing, which measures whether people find what they are looking for in a given structure.
- Content still to be invented: nothing to write on the cards, go through an affinity map on the ideas the team produces.
Description
Card sorting enters interface design with Tom Tullis's work on operating-system menus in 1985, then web information architecture with Jakob Nielsen and Darrell Sano in 1995. Donna Spencer gave it in 2009 the only book devoted entirely to the technique, the profession's reference for planning, running and interpreting a sort.
The three variants
| Variant | Who fixes the categories | Use | Its own risk |
|---|---|---|---|
| Open sort | the participants, who create the groups and name them | discovery: bringing out a structure when none has been settled | labels that vary from one participant to the next, expensive to reconcile |
| Closed sort | the researcher, who imposes the list | validation: testing a candidate tree or placing new content in it | no missing category can come to light, since the set is bounded in advance |
| Hybrid sort | the researcher, with freedom to add groups | a compromise when part of the structure is settled | the categories supplied anchor the reasoning and discourage creating a new group |
The open sort exposes the freest mental model. Hudson holds the hybrid sort to be usable while noting that it brings in the anchoring bias of the closed sort.
The closed sort keeps its use when a law or a standard imposes the classification: it then bears on the labels and on the placement within the imposed categories.
Preparing the deck
The quality of a sort is decided before the session, in the choice of what goes on the cards. A content inventory lists what exists, page by page, service by service. Duplicates and dead content are set aside, then representatives are kept: for each family suspected, a few items rather than the whole of it. A deck of several hundred cards is unworkable; a participant gives up well before the end.
How the cards are worded weighs as much as how many there are. Each card carries a short, self-contained label, in the words of the intended audience. Internal jargon has people sort the jargon rather than the content. The same word stem repeated on several cards gets them filed in the same pile without the participant having thought about their meaning, which the Nielsen Norman Group describes under the name terminology matching: grouping by resemblance of words costs less effort than grouping by the nature of the task. Two counters: vary the wording within a family and have participants say out loud why a card joined a given group.
Hudson gives orders of magnitude for duration: about 20 minutes for 30 items, 30 minutes for 50, an hour for 100. Beyond that, fatigue degrades the sort more than it adds material.
Choosing the setup
A moderated sort runs with a facilitator, in person or remotely. The participant thinks out loud; the facilitator asks why a card was placed where it was, clears up an ambiguous label, watches the hesitations. The reasons behind the classification come out. An unmoderated sort is done alone, through an online tool. It scales, reaches dispersed audiences and produces clean data, without delivering the reasons.
Tullis and Wood addressed the number of participants needed in 2004, on a set of 168 participants, by resampling subgroups and measuring the correlation between the groupings obtained and those of the full sample: about 0.90 at 15 participants, 0.93 at 20, 0.95 at 30 and 0.98 at 60. They conclude that 60 participants are a waste and that returns fall away sharply past 20 to 30. The common recommendation of 15 to 30 participants goes back to that study.
The physical setup, index cards on a table, keeps an advantage in a moderated session: handling the cards sustains engagement and makes the hesitations visible. In exchange it means transcribing every pile by hand. Digital tools record the sort and compute the matrix and the dendrogram, which makes them the ordinary setup as soon as more than a dozen participants are in view.
Running the session
The instruction fits in one sentence: put these cards into groups that belong together, your way. The participant is told that there is no right answer, that the number of groups is free and that a card which is not understood can be set aside rather than filed at random. In an open sort, the participant names the groups once the classification has settled; naming first steers the sort towards the label. Cards set aside are a result: a card nobody knows where to place signals an obscure label or content without an audience.
A frequent drift is to correct the participant or to let on that one group is more correct than another. A facilitator who prompts a category collects their own architecture. Their one legitimate intervention concerns the understanding of a word.
Analysing the sorts
The raw data are, for each participant, the list of their groups with the cards they contain and the label they were given.
The group-by-group reading (items-by-groups in the tools) brings together the labels proposed and the cards each of them attracts. It works on vocabulary: when twelve participants out of twenty-two write "Living here" and five "Residence", the users' word is known.
The pair-by-pair reading (items-by-items in the tools) counts, for each pair of cards, how many participants filed them together. That count forms a similarity matrix, which is put through agglomerative hierarchical clustering. The result is read on a dendrogram, a tree in which two cards join at a height that reflects how often they were brought together: the lower the junction, the firmer the agreement.
Hudson notes a limitation inherent in the dendrogram: each item hangs from a single branch. A card that half the panel files under A and the other half under B ends up arbitrarily on one side, and the disagreement, which is design information, disappears from the output. The counter is to keep the similarity matrix and look in it for the pairs whose score sits around half the panel.
The standard analysis delivers a single level of grouping. A hierarchy of several levels is poorly served by the common tools and is better handled when part of the tree is fixed in advance.
The cards are sorted, finally, out of context, with no page around them, no visual hierarchy, no search engine. The groups obtained say how people classify, which does not always coincide with how they would navigate the product in service.
AI considerations
The most profitable help sits upstream of the session. A language model deduplicates an inventory of several hundred pages, proposes a small number of representatives for each family suspected and flags the labels whose word stem repeats. On a multilingual Swiss site, it normalises the French, German and Italian wordings so that one deck serves three comparable panels. Downstream, it codes the transcripts of moderated sessions and brings synonymous labels together.
One thing is forbidden to it: sorting the cards in place of the participants. A model asked to produce the grouping returns the implicit structure of its training corpus, when the technique exists to measure that of the intended audience. The same caution applies to the category labels a model suggests: they are plausible because they are frequent elsewhere, which is the opposite of the local vocabulary being sought.
Transcripts of moderated sessions carry identifiable statements and sometimes personal data. Entrusting them to an external service engages Swiss data-protection obligations under the revised FADP: an internal or approved tool, otherwise anonymisation before any automated processing.
Examples
A commune in French-speaking Switzerland is redesigning the navigation of its website: the 62 online services listed in the content inventory were put through an open, unmoderated sort by 22 residents recruited across three age brackets.
| Group named by the participants | Cards most often brought together | Share of the panel |
|---|---|---|
| Living here and moving | registration on arrival, change of address, certificate of residence, parking permit | 19 / 22 |
| Building and renovating | building permit application, public inspection period, connection charge, rubble disposal | 18 / 22 |
| Family and school | enrolment in compulsory school, after-school care, family allowances, holiday camp | 16 / 22 |
| Taxes and charges | tax instalment, refuse-bag charge, water bill, visitor's tax | 14 / 22 |
| No stable group | registering a dog: filed under "Living here and moving" by 11 residents, under "Taxes and charges" by 9 | split 11 / 9 |
The first four groups read as candidate top-level sections. The pair-by-pair reading of the same panel puts numbers on how solid they are: certificate of residence and change of address were filed together by 21 residents out of 22, enrolment in compulsory school and after-school care by 20. At that level the section is settled and the discussion is about its label.
The contested card is read on the same matrix. Registering a dog and tax instalment score 9 out of 22; registering a dog and certificate of residence, 11 out of 22. Two scores close to half the panel signal a split that the dendrogram would settle without saying so. What follows is a design decision rather than a statistical arbitration: place the service in both spots or file it on one side and make it findable from the other through a cross-reference.
Visualisations
In the raw sorts, look for the cards a participant set aside and those they moved several times. In the similarity matrix, go first to the scores close to half the panel, which mark the splits. In the dendrogram, read the height of the junctions before the branches: it is what says whether a group holds. The matrix stays open next to the dendrogram throughout the arbitration, since it is the only one of the three that keeps the disagreements.
Analysis chain of a card sort
from raw sorts to the dendrogram
C and D: 1 out of 3, a split between two branches of the dendrogram.
Cost
| Phase | Level | Justification |
|---|---|---|
| Preparation | High | Content inventory, selection of representatives, writing the labels, recruiting the panel. The sort takes an hour, choosing what goes on the cards takes days. |
| Execution | Low to medium | Low in an unmoderated sort, where the session runs without intervention; medium in a moderated sort, where each session occupies a facilitator for 20 to 60 minutes depending on the number of cards. |
| Documentation | Medium | Similarity matrix, dendrogram, proposed tree with the departures from the sort justified. An online tool produces the first two automatically. |
Tooling
The minimal setup is a deck of index cards, a marker and a table, plus a spreadsheet to enter the piles and compute the co-occurrences. It suits a handful of moderated sessions in person; manual transcription sets its limit.
The specialised platforms, Optimal Workshop's OptimalSort, UXtweak, Maze, kardSort, cover the three variants, host the unmoderated sort and deliver the similarity matrix and the dendrogram with no further processing. A digital whiteboard, Miro or Mural, will do for a moderated remote sort but analyses nothing: the piles have to be entered afterwards.
For a Swiss organisation, where the data are hosted steers the choice when the panel is identifiable or when the content being sorted is not public. An unmoderated sort on a tool hosted outside Switzerland is run with pseudonymised participants. The analysis can also be redone by hand: agglomerative hierarchical clustering is computed in R or in Python from the exported similarity matrix alone.
Sources
- Hudson, W., "Card Sorting", in The Encyclopedia of Human-Computer Interaction, 2nd ed., Interaction Design Foundation, 2014: documented history of the method, definition of the three variants, analysis and limitations. Chapter online.
- Spencer, D., Card Sorting: Designing Usable Categories, Rosenfeld Media, 2009: the only book devoted entirely to the technique, the profession's reference for planning, running and interpreting a sort. Publisher's page.
- Tullis, T. and Wood, L., "How Many Users Are Enough for a Card-Sorting Study?", UPA 2004 Conference, Minneapolis, 2004 (Usability Professionals' Association): the study the participant numbers come from.
- Nielsen Norman Group, "Card Sorting: Uncover Users' Mental Models for Better Information Architecture": definition and framing of the technique.
- Nielsen Norman Group, "Card Sorting: Pushing Users Beyond Terminology Matches": grouping by resemblance of words and the facilitator's counters.

