L
🟢 Information Architecture · Level 3 · Labels

The Words on the Doors

A label is a promise about what is behind it. Keep the promise and browsing feels nearly effortless; break it once and the reader starts checking every neighbouring word too. This is the least glamorous part of the discipline and the highest-leverage — a correct first click predicts task success at around 87%, a wrong one at about 46%1. In a text-only structure, most of that opening decision is carried by the words — position, grouping and prior habit carry the rest.

Predict · before the clickPromise · kept on arrivalProve · first click, real task25 questions · 7 formats

What’s going on

A label is a contract

Every word in a menu makes a small promise: click here and you will get roughly this. Readers use those promises to move without thinking — they predict, they click, and the destination either confirms the prediction or teaches them the label cannot be trusted. That second outcome costs far more than a single wasted click. After one broken promise people stop browsing confidently and start checking defensively, opening things to see what is inside rather than reading their way there. The cheap behaviour collapses into an expensive one, and it happens across the whole menu, not only at the word that lied.

Why the first click carries so much

Bailey and Wolfson measured what the opening move is worth. When someone’s first click is correct, they finish the task around 87% of the time. When it is wrong, that falls to roughly 46%1. Half the people who start down the wrong path never come back, even though going back costs one click. Confidence is what actually gets spent, and a wrong turn spends a lot of it. The useful consequence is practical: first-click testing needs one screen, one task and no build at all, and it measures the labels with everything else held out of the way.

Where it goes wrong
  • The label is written from an idea of the section → rather than from a list of what is actually in it, so it drifts out of date the moment the contents change.
  • Brevity is the only test applied → short words fit the grid beautifully and predict nothing. “More” fits perfectly.
  • The team reviews its own menu → everyone in the room can decode “Revenues” instantly, which is exactly what disqualifies them as testers.

Click first, then find out

Four tasks. Five labels each. One click, no going back — because in the research that is the click that matters. Where you land is not scored on cleverness; it is scored on whether the words told you the truth.

🖱️ The first-click test

Your task
Right first click
0
Tasks done
0 / 4
If right
~87%
If wrong
~46%

The two figures on the right are the research, not your score: they are the probability of finishing the task given a right or wrong opening move. Notice how often the wrong item is the reasonable one — a label that could plausibly hold the answer is more dangerous than one that obviously could not.

The four tests

NN/g’s four Ss give a label four separate ways to fail2. They pull against each other on purpose: specific and substantial push a label longer, succinct pushes back, and sincere decides whether the promise can be kept at all. Teams usually optimise the last one alone, because it is the only one you can judge by eye.

🏷️ Score the label

The destination
Five candidate labels for that one page. Pick each in turn and see which tests it passes.

Two vocabularies, one door

A hospital corridor says Radiology. The patient is looking for an X-ray. Both words are correct — one is correct for the institution and one for the person walking towards it. The failure is picking a side and pretending the other vocabulary does not exist.

🗣️ Say it the other way

Here is the word the organisation uses. What would the visitor be looking for?
Matched
0
Seen
0 / 8

The usual repair keeps both words. Hospital signage has done this for decades — the department name with the plain words underneath — because the referral letter, the staff rota and the frightened visitor all have a legitimate claim on the same door.

Six kinds of label

The four tests were written for link and navigation text. Applying them identically to a status chip or a destructive button quietly misfires — these six do different jobs, and only the first is really about predicting a destination.

Navigation labelPredict a destinationCouncil tax
Action labelPredict a consequenceDelete permanently
Status labelDescribe the present stateAwaiting approval
Category labelDescribe what members shareBins and recycling
Field labelIdentify what is being asked forPostcode
InstructionExplain how to proceedEnter the code we texted you

A navigation label predicts where you will land; an action label predicts what will happen to your data. That difference is why “Delete” and “Reports” cannot be judged by the same rulebook, and why the section below exists.

What breaks a label

Six failures, six different repairs
  • Invented words → nobody searches for a term they have never met, so a coined name is invisible while browsing and invisible in search. Familiar words beat invented ones — though an established domain term beats a vague simplification wherever the audience actually holds it.
  • “Learn more” → fails substantial. Lifted out of its sentence it names nothing, which is how a scanner meets it, and how a screen-reader user may meet it when moving through a page’s links. Replace it with the destination’s own name.
  • “More” → a decision that was never made, wearing the costume of a category. Read its contents aloud: if a shared property emerges, name it; if none does, you have found a gap in the structure.
  • One word, two jobs → “Orders” covering both placing one and reviewing past ones. Split it: Track an order and Order again.
  • Load-bearing endings → truncation eats backwards, so anything that only becomes distinguishable in the last two words vanishes exactly where space is tightest.
  • Length fitted to one language → teams often budget roughly 30% expansion for German, and short interface strings can grow far more than that. A horizontal bar tuned to English word lengths has no slack; a vertical list simply gets taller.
Front-load the difference

Read three sibling labels that open with the same two words — Guide to billing, Guide to invoices, Guide to refunds — and notice how much of the scan is spent on the part they share. People read the openings of items in a list, so whatever separates an entry from its neighbours should arrive first and whatever it holds in common should follow. This is the same argument as truncation seen from the other end, and it costs nothing: the words are already there, in the wrong order.

The fifth test: does it separate?

The four Ss judge a label alone. A menu is a vocabulary system, and a word can pass all four and still fail beside its neighbours. Add a contextual test after the individual ones — separating: does this label clearly distinguish itself from the labels next to it?

🔍 The collision detector

Three real sibling sets. Read each one and decide what is wrong before revealing it.
Hold one grammar, and choose the part of speech deliberately

Sibling labels should usually share a shape. A menu reading Payments · Track your order · Returns policy · Get help · My profile changes grammatical form on every line, and the reader pays for the switching even when no single item is ambiguous. Pick a system and hold it. Beyond consistency there is a choice worth making on purpose: nouns name stable spaces and objects, verbs name actions, and states name a present condition. “Projects” beats “Manage projects” when the destination is a place you go; “Create project” is right when the control does something; “Draft” is right when the word describes what a thing currently is. Mixing the three inside one row is how a menu ends up feeling vaguely wrong to everybody and clearly wrong to nobody.

In the real world

Hospital signage

A corridor sign reads Radiology, and underneath, in the same size, X-rays and scans. Neither word is wrong and neither is deleted. The referral letter, the staff rota and the frightened visitor all have a legitimate claim on that door, so the door carries both names. Decades of wayfinding practice arrived at pairing rather than replacement.

The cloud console

Delete, remove, detach, deactivate, archive, terminate. Six words that sound like synonyms and trigger six different outcomes — one hides, two are reversible, one breaks a relationship, one destroys attached storage. When consequences differ, precision stops being style and becomes safety: the label has to name what you cannot undo.

Netflix's 76,897 altgenres

Sometimes a label fails because the category underneath it is too broad to name honestly. “Drama” had stopped separating anything, so Netflix went finer rather than simpler — tens of thousands of micro-categories tagged against a 36-page manual, with romantic appearing in 5,272 of them. The word was not the problem; the grouping was.

When the word is the consequence

Navigation labels cost a wasted click. Action labels cost data. Six words that sound like synonyms, ordered by what they actually do — guess where each one sits before you reveal it.

⚠️ The destructive-verb lab

Click a word to see what it does in a typical cloud console. The bar shows how much of it you cannot take back.

Two of these are reversible, one hides without destroying, one breaks a relationship between objects, and one takes attached storage with it. “Remove” is the worst of the set precisely because it says nothing about which — the reader has to already know the system to predict the outcome.

And the same object, named four ways
On the websiteOrder number
In the confirmation emailReference
On the support lineCase ID
In the courier's appTracking number

One parcel, four vocabularies, four teams. The customer holding a reference from an email cannot tell whether the support agent asking for a case ID wants the same string. A controlled vocabulary — one preferred term per concept, written down, with its forbidden synonyms beside it — is unglamorous governance that prevents exactly this.

Name the menu

Six sections, every one named after whoever owns it internally. You have the contents. Name each from what is actually inside — then the page runs a first-click check against the labels you chose. Three sectors in the bank: council, hospital, warehouse.

🏛️ The council menu

Test yourself — a mixed set

Seven question formats, the way Beyond Dictionary serves them. Every question carries layered hints — a nudge, the reasoning, then a deeper connection — so a wrong answer opens a door instead of closing one.

Question 1 of 25
MCQ

The words, beyond the basics

The follow-on questions — the ones that arrive once everyone agrees labels matter and the argument moves to which words.

What makes a navigation label good?
ConceptualWhatcomplexity 2

Two things: a reader can predict what is behind it before clicking, and the destination then keeps that promise. NN/g's four tests name the ways it fails: specific (does it name the real destination), sincere (does the page deliver what was promised), substantial (does it still mean something lifted out of its sentence), and succinct (is it no longer than clarity requires). The four pull against each other deliberately — and there is no maximum word count, because concision is subordinate to the other three rather than above them.

Why does a user's first click matter more than the clicks that follow it?
AnalyticalWhycomplexity 3

Because it predicts the outcome. Bailey and Wolfson found that a correct first click is followed by task success around 87% of the time, while a wrong first click drops that to roughly 46%. Going back costs one click, so what people are actually losing is confidence rather than time. The practical consequence is that first-click testing is unusually cheap evidence — one screen, one task, no build — and it isolates the labels from everything else in the experience.

If an organisation's department names are accurate, why should its website menu not use them?
ApplicationHowcomplexity 3

Because accuracy for the institution is not the same as usefulness for the visitor. "Bursary", "Revenues" and "Development Management" are all correct names, and every one of them requires the reader to already know how the organisation is arranged. Task names — pay your fees, council tax, planning and building — require nothing. Keep the department names where they belong, on the pages and in correspondence, and let the menu name the errand. Hospital signage has run both for decades: the department name with the plain words beneath it.

What is wrong with "Learn more"?
ScenarioWhatcomplexity 3

The substantial test is the one it fails. People scan for links rather than reading the sentences around them, and a screen reader can list every link on a page — where a column of identical "Learn more" entries names nothing at all. The repair costs nothing: replace the empty phrase with the destination's own name, so "Learn more" becomes "How refunds work". The surrounding sentence almost always reads better afterwards, because it no longer has to carry the meaning the link was supposed to.

How can a disagreement about navigation labels be settled with evidence rather than taste?
AnalyticalHowcomplexity 4

Convert the opinion into a testable claim. Rather than debating whether "Solutions" is clear, run a first-click test: three or four real tasks, the current label against two contents-derived alternatives, with people who do not work on the product. It takes an afternoon and produces behaviour rather than preference — which matters, because people routinely prefer a label they then fail to use. It also changes the politics: nobody has to be wrong in a meeting, because the test is what turns out to be right.

Is a taxonomy with thousands of narrow categories ever better than one with a few broad ones?
ReflectiveWhencomplexity 4

Yes, when the problem is granularity rather than ambiguity. Film genre stopped discriminating once "Drama" held tens of thousands of unlike titles, and Netflix's response was to go finer rather than simpler — roughly 76,897 micro-categories, tagged by people working from a 36-page manual. Compare that with a supermarket, which handles coconut milk belonging in three aisles by stocking it in several places rather than inventing new categories. Both are correct responses to different faults: one vocabulary was too coarse, the other genuinely ambiguous.

Why do teams keep shipping navigation labels they know are vague?
ReflectiveWhycomplexity 5

Because vagueness is safe internally and expensive only externally. A container word like "Solutions" or "More" never has to be argued about, accommodates every team's content, and fits the grid neatly. The cost lands somewhere nobody in the room is sitting — on a reader deciding whether it is worth a click. That asymmetry is why label decisions need evidence rather than consensus: the people in the meeting are the only ones who already know what is behind the word, which is precisely what disqualifies their judgement.

Glossary — the 12 words that name it

Label

What it means
The word on a door: a promise about what is behind it.
Why it matters
Readers predict from it and either have the prediction confirmed or learn not to trust the menu.
Key question
Could someone say what is behind this without clicking?

First-click testing

What it means
One screen, one task, one click — measuring where the hand goes first.
Why it matters
A correct first click predicts success at about 87%, a wrong one at about 46%.
Key question
Where does the very first click land, and where does it go wrong?

Specific

What it means
The label names the actual destination rather than a category it might belong to.
Why it matters
Vague words make every click a gamble, which turns cheap browsing into defensive checking.
Key question
Does this name the thing, or a container the thing might be in?

Sincere

What it means
The destination delivers what the label promised.
Why it matters
The only one of the four tests you cannot judge without following the link.
Key question
Did the page keep the promise the word made?

Substantial

What it means
The label still means something when lifted out of its sentence.
Why it matters
Scanners and screen readers meet links as a list, with no surrounding prose to lean on.
Key question
Read aloud on its own, does this name anything?

Succinct

What it means
No longer than clarity requires — and subordinate to the other three.
Why it matters
There is no maximum word count; brevity is the constraint, not the goal.
Key question
Is every word here earning its place?

Expert and everyday vocabulary

What it means
The same thing named for the institution and named for the visitor.
Why it matters
Both have a legitimate claim, so the durable answer usually carries both rather than choosing.
Key question
Which of these two words did the reader arrive already holding?

Polysemy

What it means
One word carrying two or more distinct meanings.
Why it matters
The quietest failure, because the word looks perfectly clear to whoever wrote it.
Key question
Does this label mean two things if you say it twice?

Open card sorting

What it means
Asking people to group content and then name the groups themselves.
Why it matters
It surfaces vocabulary nobody in the room would have proposed.
Key question
What did they call it when nobody offered them a word?

Controlled vocabulary

What it means
An agreed set of terms used consistently across a product.
Why it matters
Without one, the same thing acquires three names and continuity breaks between devices and teams.
Key question
Is this thing called the same thing everywhere it appears?

Front-loading

What it means
Putting the distinguishing word at the start of a label.
Why it matters
Scanning reads openings and truncation eats endings — both argue for the same order.
Key question
If this were cut in half, would the remaining half still identify it?

Container word

What it means
A label broad enough to hold anything: Solutions, Resources, More, Explore.
Why it matters
Safe to agree on internally and impossible to predict from outside.
Key question
Could this word plausibly hold three completely different sections?

The evidence behind this lesson

Every number quoted above, with where it comes from and why it is here.

1 · 87% and 46%
Bob Bailey and Cari Wolfson’s analysis of scenario-based usability tests: a correct first click is followed by task success around 87% of the time, a wrong one around 46%. Summarised with the method at Optimal Workshop — correct first clicks and task success, with a look at how well click tests predict live behaviour at MeasuringU.
2 · The four Ss
Specific, sincere, substantial, succinct — and the point that concision is subordinate to the other three, with no maximum word count: NN/g — Better Link Labels.
76,897 altgenres
How Netflix answered a vocabulary that had stopped discriminating — human tagging against a 36-page manual, with “romantic” appearing in 5,272 categories: FlowingData and NPR.
Plain words, in public service
The style guidance behind “fill in your tax return” over “self-assessment”, including a standing list of words to avoid: the GOV.UK style guide.
The canon
Rosenfeld, Morville & Arango, Information Architecture for the Web and Beyond; Abby Covert’s How to Make Sense of Any Mess, which treats language as the material itself; and Steve Krug’s Don’t Make Me Think for the argument that every moment of hesitation is a cost.
The full reference shelf
Every book, study, standard and open argument behind all four parts is gathered in one place at the end of Part 4: Where this comes from.

You can name it now

Four things changed in how you read a menu — and one place to take them next.

Prediction
You read a label as a promise now, and notice immediately when it cannot be kept.
Evidence
You can settle a naming argument in an afternoon instead of a meeting.
Vocabulary
You hear two legitimate words for the same door, and stop choosing between them.
Diagnosis
Container words stand out, and you know which repair each failure needs.
Up next · Part 4

When It Breaks

Structures fail in patterns — and the same levers that help people can be turned against them. The ethics chapter.

Continue to Part 4 →