Your Training Partner
Techniques Toolbox
Analysis chain of a card sort: the participants' raw sorts, the similarity matrix counting the card pairs filed together, then the dendrogram whose junction height marks the strength of the agreement.

Card Sorting

Card sorting is a user-research method. Participants are given a deck of cards, each carrying an existing piece of content, a page, a service, a form, then they group them the way that makes sense to them and name the groups. Those groups and their labels expose the mental model by which these people expect to find the material. The result feeds an information architecture: a navigation menu, a taxonomy, the outline of a document. A sort produces a candidate structure, which is then put through a separate check.

Goal

In a card sort, users group existing pieces of content into the categories that feel natural to them, then they name those categories. The method answers a design question: how does the audience expect to find this material, when the vocabulary and the divisions of the organisation that publishes it are not imposed on them?

It supports the choice of an information architecture: the top-level sections of a site, the order of a menu, the entries of a document taxonomy. The deliverable has two halves: on one side the groups as the participants formed and labelled them, qualitative material about vocabulary; on the other the count of card pairs filed together, quantitative material about the strength of each association. The two are read together: when a firm association attracts diverging labels, the group holds and the work is on the label.

Measuring whether people find what they are looking for in the resulting structure calls for tree testing.

Card sorting and the affinity map

The two techniques look alike from a distance, cards moved by hand until groups appear; the direction of the exercise separates them.

Who sortsWhat is sortedWhat comes out
Card sortingusers or stakeholders from outside the teampieces of content that already existthe organisation those people see in it, the basis of an information architecture
Affinity mapthe team itselfthe ideas it has just produced in sessionconvergence on a few themes after a phase of divergence

If the material to be sorted was written within the hour by the people in the room, the affinity map is the tool; if it already exists and the question is how an audience files it, this is a card sort, with participants recruited from outside the team.

Card sorting settles the grouping. Settling the meaning and the vocabulary of the terms belongs to concept modelling and to the glossary; eliciting one by one the criteria by which a person tells two items apart belongs to the repertory grid.

Usage

When to use it

  • Redesign of an existing navigation: the content inventory is available and the current tree is contested.
  • Internal vocabulary far from the public's: the site is divided along the org chart and visitors do not recognise the labels.
  • A candidate taxonomy to put to the test: a closed sort, whose categories are fixed in advance, shows which of them take content without hesitation.
  • Internal disagreement about classification: the sort moves the arbitration from departmental opinion to user data.
  • A dispersed audience: an unmoderated remote sort recruits widely at a low marginal cost per participant.

When not to use it

  • A structure to evaluate rather than to produce: use tree testing, which measures whether people find what they are looking for in a given structure.
  • Content still to be invented: nothing to write on the cards, go through an affinity map on the ideas the team produces.

Description

Card sorting enters interface design with Tom Tullis's work on operating-system menus in 1985, then web information architecture with Jakob Nielsen and Darrell Sano in 1995. Donna Spencer gave it in 2009 the only book devoted entirely to the technique, the profession's reference for planning, running and interpreting a sort.

The three variants

VariantWho fixes the categoriesUseIts own risk
Open sortthe participants, who create the groups and name themdiscovery: bringing out a structure when none has been settledlabels that vary from one participant to the next, expensive to reconcile
Closed sortthe researcher, who imposes the listvalidation: testing a candidate tree or placing new content in itno missing category can come to light, since the set is bounded in advance
Hybrid sortthe researcher, with freedom to add groupsa compromise when part of the structure is settledthe categories supplied anchor the reasoning and discourage creating a new group

The open sort exposes the freest mental model. Hudson holds the hybrid sort to be usable while noting that it brings in the anchoring bias of the closed sort.

The closed sort keeps its use when a law or a standard imposes the classification: it then bears on the labels and on the placement within the imposed categories.

Preparing the deck

The quality of a sort is decided before the session, in the choice of what goes on the cards. A content inventory lists what exists, page by page, service by service. Duplicates and dead content are set aside, then representatives are kept: for each family suspected, a few items rather than the whole of it. A deck of several hundred cards is unworkable; a participant gives up well before the end.

How the cards are worded weighs as much as how many there are. Each card carries a short, self-contained label, in the words of the intended audience. Internal jargon has people sort the jargon rather than the content. The same word stem repeated on several cards gets them filed in the same pile without the participant having thought about their meaning, which the Nielsen Norman Group describes under the name terminology matching: grouping by resemblance of words costs less effort than grouping by the nature of the task. Two counters: vary the wording within a family and have participants say out loud why a card joined a given group.

Hudson gives orders of magnitude for duration: about 20 minutes for 30 items, 30 minutes for 50, an hour for 100. Beyond that, fatigue degrades the sort more than it adds material.

Choosing the setup

A moderated sort runs with a facilitator, in person or remotely. The participant thinks out loud; the facilitator asks why a card was placed where it was, clears up an ambiguous label, watches the hesitations. The reasons behind the classification come out. An unmoderated sort is done alone, through an online tool. It scales, reaches dispersed audiences and produces clean data, without delivering the reasons.

Tullis and Wood addressed the number of participants needed in 2004, on a set of 168 participants, by resampling subgroups and measuring the correlation between the groupings obtained and those of the full sample: about 0.90 at 15 participants, 0.93 at 20, 0.95 at 30 and 0.98 at 60. They conclude that 60 participants are a waste and that returns fall away sharply past 20 to 30. The common recommendation of 15 to 30 participants goes back to that study.

The physical setup, index cards on a table, keeps an advantage in a moderated session: handling the cards sustains engagement and makes the hesitations visible. In exchange it means transcribing every pile by hand. Digital tools record the sort and compute the matrix and the dendrogram, which makes them the ordinary setup as soon as more than a dozen participants are in view.

Running the session

The instruction fits in one sentence: put these cards into groups that belong together, your way. The participant is told that there is no right answer, that the number of groups is free and that a card which is not understood can be set aside rather than filed at random. In an open sort, the participant names the groups once the classification has settled; naming first steers the sort towards the label. Cards set aside are a result: a card nobody knows where to place signals an obscure label or content without an audience.

A frequent drift is to correct the participant or to let on that one group is more correct than another. A facilitator who prompts a category collects their own architecture. Their one legitimate intervention concerns the understanding of a word.

Analysing the sorts

The raw data are, for each participant, the list of their groups with the cards they contain and the label they were given.

The group-by-group reading (items-by-groups in the tools) brings together the labels proposed and the cards each of them attracts. It works on vocabulary: when twelve participants out of twenty-two write "Living here" and five "Residence", the users' word is known.

The pair-by-pair reading (items-by-items in the tools) counts, for each pair of cards, how many participants filed them together. That count forms a similarity matrix, which is put through agglomerative hierarchical clustering. The result is read on a dendrogram, a tree in which two cards join at a height that reflects how often they were brought together: the lower the junction, the firmer the agreement.

Hudson notes a limitation inherent in the dendrogram: each item hangs from a single branch. A card that half the panel files under A and the other half under B ends up arbitrarily on one side, and the disagreement, which is design information, disappears from the output. The counter is to keep the similarity matrix and look in it for the pairs whose score sits around half the panel.

The standard analysis delivers a single level of grouping. A hierarchy of several levels is poorly served by the common tools and is better handled when part of the tree is fixed in advance.

The cards are sorted, finally, out of context, with no page around them, no visual hierarchy, no search engine. The groups obtained say how people classify, which does not always coincide with how they would navigate the product in service.

AI considerations

The most profitable help sits upstream of the session. A language model deduplicates an inventory of several hundred pages, proposes a small number of representatives for each family suspected and flags the labels whose word stem repeats. On a multilingual Swiss site, it normalises the French, German and Italian wordings so that one deck serves three comparable panels. Downstream, it codes the transcripts of moderated sessions and brings synonymous labels together.

One thing is forbidden to it: sorting the cards in place of the participants. A model asked to produce the grouping returns the implicit structure of its training corpus, when the technique exists to measure that of the intended audience. The same caution applies to the category labels a model suggests: they are plausible because they are frequent elsewhere, which is the opposite of the local vocabulary being sought.

Transcripts of moderated sessions carry identifiable statements and sometimes personal data. Entrusting them to an external service engages Swiss data-protection obligations under the revised FADP: an internal or approved tool, otherwise anonymisation before any automated processing.

Examples

A commune in French-speaking Switzerland is redesigning the navigation of its website: the 62 online services listed in the content inventory were put through an open, unmoderated sort by 22 residents recruited across three age brackets.

Group named by the participantsCards most often brought togetherShare of the panel
Living here and movingregistration on arrival, change of address, certificate of residence, parking permit19 / 22
Building and renovatingbuilding permit application, public inspection period, connection charge, rubble disposal18 / 22
Family and schoolenrolment in compulsory school, after-school care, family allowances, holiday camp16 / 22
Taxes and chargestax instalment, refuse-bag charge, water bill, visitor's tax14 / 22
No stable groupregistering a dog: filed under "Living here and moving" by 11 residents, under "Taxes and charges" by 9split 11 / 9
Group-by-group reading of an open sort: the labels come from the participants. The last row carries the most useful result: registering a dog splits the panel in two, a card that a dendrogram would attach by default to a single branch.

The first four groups read as candidate top-level sections. The pair-by-pair reading of the same panel puts numbers on how solid they are: certificate of residence and change of address were filed together by 21 residents out of 22, enrolment in compulsory school and after-school care by 20. At that level the section is settled and the discussion is about its label.

The contested card is read on the same matrix. Registering a dog and tax instalment score 9 out of 22; registering a dog and certificate of residence, 11 out of 22. Two scores close to half the panel signal a split that the dendrogram would settle without saying so. What follows is a design decision rather than a statistical arbitration: place the service in both spots or file it on one side and make it findable from the other through a cross-reference.

Visualisations

In the raw sorts, look for the cards a participant set aside and those they moved several times. In the similarity matrix, go first to the scores close to half the panel, which mark the splits. In the dendrogram, read the height of the junctions before the branches: it is what says whether a group holds. The matrix stays open next to the dendrogram throughout the arbitration, since it is the only one of the three that keeps the disagreements.

Analysis chain of a card sort

from raw sorts to the dendrogram

Raw sorts
3 participants, their piles
P1
ABC
DE
F
P2
ABC
DEF
P3
AB
CDE
F
Similarity matrix
card pairs filed together, out of 3
A
B
C
D
E
F
A
3
2
0
0
0
B
3
2
0
0
0
C
2
2
1
1
0
D
0
0
1
3
1
E
0
0
1
3
1
F
0
0
0
1
1
0 1 2 3 participants out of 3

C and D: 1 out of 3, a split between two branches of the dendrogram.

Dendrogram
the height of the junctions, not the branches
Schematic dendrogram of six cards A to FA and B join at the lowest height, as do D and E: their agreement is the firmest. C joins the A-B group a little higher, F joins the D-E group higher still, then the two branches join at the very top, at the weakest height.frequency of the pairinglow junction:firm agreementhigh junction: weak associationABCDEF
The three representations of the analysis chain. On the dendrogram, the height of the junctions measures how often two cards were paired, on no fixed scale: a low junction marks firm agreement, a high junction a weak association.

Cost

PhaseLevelJustification
PreparationHighContent inventory, selection of representatives, writing the labels, recruiting the panel. The sort takes an hour, choosing what goes on the cards takes days.
ExecutionLow to mediumLow in an unmoderated sort, where the session runs without intervention; medium in a moderated sort, where each session occupies a facilitator for 20 to 60 minutes depending on the number of cards.
DocumentationMediumSimilarity matrix, dendrogram, proposed tree with the departures from the sort justified. An online tool produces the first two automatically.

Tooling

The minimal setup is a deck of index cards, a marker and a table, plus a spreadsheet to enter the piles and compute the co-occurrences. It suits a handful of moderated sessions in person; manual transcription sets its limit.

The specialised platforms, Optimal Workshop's OptimalSort, UXtweak, Maze, kardSort, cover the three variants, host the unmoderated sort and deliver the similarity matrix and the dendrogram with no further processing. A digital whiteboard, Miro or Mural, will do for a moderated remote sort but analyses nothing: the piles have to be entered afterwards.

For a Swiss organisation, where the data are hosted steers the choice when the panel is identifiable or when the content being sorted is not public. An unmoderated sort on a tool hosted outside Switzerland is run with pseudonymised participants. The analysis can also be redone by hand: agglomerative hierarchical clustering is computed in R or in Python from the exported similarity matrix alone.

Sources

  • Hudson, W., "Card Sorting", in The Encyclopedia of Human-Computer Interaction, 2nd ed., Interaction Design Foundation, 2014: documented history of the method, definition of the three variants, analysis and limitations. Chapter online.
  • Spencer, D., Card Sorting: Designing Usable Categories, Rosenfeld Media, 2009: the only book devoted entirely to the technique, the profession's reference for planning, running and interpreting a sort. Publisher's page.
  • Tullis, T. and Wood, L., "How Many Users Are Enough for a Card-Sorting Study?", UPA 2004 Conference, Minneapolis, 2004 (Usability Professionals' Association): the study the participant numbers come from.
  • Nielsen Norman Group, "Card Sorting: Uncover Users' Mental Models for Better Information Architecture": definition and framing of the technique.
  • Nielsen Norman Group, "Card Sorting: Pushing Users Beyond Terminology Matches": grouping by resemblance of words and the facilitator's counters.
Business Visualizations
All techniques
CATWOE Analysis