A label is a promise about what is behind it. Keep the promise and browsing feels nearly effortless; break it once and the reader starts checking every neighbouring word too. This is the least glamorous part of the discipline and the highest-leverage — a correct first click predicts task success at around 87%, a wrong one at about 46%1. In a text-only structure, most of that opening decision is carried by the words — position, grouping and prior habit carry the rest.
Every word in a menu makes a small promise: click here and you will get roughly this. Readers use those promises to move without thinking — they predict, they click, and the destination either confirms the prediction or teaches them the label cannot be trusted. That second outcome costs far more than a single wasted click. After one broken promise people stop browsing confidently and start checking defensively, opening things to see what is inside rather than reading their way there. The cheap behaviour collapses into an expensive one, and it happens across the whole menu, not only at the word that lied.
Bailey and Wolfson measured what the opening move is worth. When someone’s first click is correct, they finish the task around 87% of the time. When it is wrong, that falls to roughly 46%1. Half the people who start down the wrong path never come back, even though going back costs one click. Confidence is what actually gets spent, and a wrong turn spends a lot of it. The useful consequence is practical: first-click testing needs one screen, one task and no build at all, and it measures the labels with everything else held out of the way.
Four tasks. Five labels each. One click, no going back — because in the research that is the click that matters. Where you land is not scored on cleverness; it is scored on whether the words told you the truth.
The two figures on the right are the research, not your score: they are the probability of finishing the task given a right or wrong opening move. Notice how often the wrong item is the reasonable one — a label that could plausibly hold the answer is more dangerous than one that obviously could not.
NN/g’s four Ss give a label four separate ways to fail2. They pull against each other on purpose: specific and substantial push a label longer, succinct pushes back, and sincere decides whether the promise can be kept at all. Teams usually optimise the last one alone, because it is the only one you can judge by eye.
A hospital corridor says Radiology. The patient is looking for an X-ray. Both words are correct — one is correct for the institution and one for the person walking towards it. The failure is picking a side and pretending the other vocabulary does not exist.
The usual repair keeps both words. Hospital signage has done this for decades — the department name with the plain words underneath — because the referral letter, the staff rota and the frightened visitor all have a legitimate claim on the same door.
The four tests were written for link and navigation text. Applying them identically to a status chip or a destructive button quietly misfires — these six do different jobs, and only the first is really about predicting a destination.
A navigation label predicts where you will land; an action label predicts what will happen to your data. That difference is why “Delete” and “Reports” cannot be judged by the same rulebook, and why the section below exists.
Read three sibling labels that open with the same two words — Guide to billing, Guide to invoices, Guide to refunds — and notice how much of the scan is spent on the part they share. People read the openings of items in a list, so whatever separates an entry from its neighbours should arrive first and whatever it holds in common should follow. This is the same argument as truncation seen from the other end, and it costs nothing: the words are already there, in the wrong order.
The four Ss judge a label alone. A menu is a vocabulary system, and a word can pass all four and still fail beside its neighbours. Add a contextual test after the individual ones — separating: does this label clearly distinguish itself from the labels next to it?
Sibling labels should usually share a shape. A menu reading Payments · Track your order · Returns policy · Get help · My profile changes grammatical form on every line, and the reader pays for the switching even when no single item is ambiguous. Pick a system and hold it. Beyond consistency there is a choice worth making on purpose: nouns name stable spaces and objects, verbs name actions, and states name a present condition. “Projects” beats “Manage projects” when the destination is a place you go; “Create project” is right when the control does something; “Draft” is right when the word describes what a thing currently is. Mixing the three inside one row is how a menu ends up feeling vaguely wrong to everybody and clearly wrong to nobody.
A corridor sign reads Radiology, and underneath, in the same size, X-rays and scans. Neither word is wrong and neither is deleted. The referral letter, the staff rota and the frightened visitor all have a legitimate claim on that door, so the door carries both names. Decades of wayfinding practice arrived at pairing rather than replacement.
Delete, remove, detach, deactivate, archive, terminate. Six words that sound like synonyms and trigger six different outcomes — one hides, two are reversible, one breaks a relationship, one destroys attached storage. When consequences differ, precision stops being style and becomes safety: the label has to name what you cannot undo.
Sometimes a label fails because the category underneath it is too broad to name honestly. “Drama” had stopped separating anything, so Netflix went finer rather than simpler — tens of thousands of micro-categories tagged against a 36-page manual, with romantic appearing in 5,272 of them. The word was not the problem; the grouping was.
Navigation labels cost a wasted click. Action labels cost data. Six words that sound like synonyms, ordered by what they actually do — guess where each one sits before you reveal it.
Two of these are reversible, one hides without destroying, one breaks a relationship between objects, and one takes attached storage with it. “Remove” is the worst of the set precisely because it says nothing about which — the reader has to already know the system to predict the outcome.
One parcel, four vocabularies, four teams. The customer holding a reference from an email cannot tell whether the support agent asking for a case ID wants the same string. A controlled vocabulary — one preferred term per concept, written down, with its forbidden synonyms beside it — is unglamorous governance that prevents exactly this.
Six sections, every one named after whoever owns it internally. You have the contents. Name each from what is actually inside — then the page runs a first-click check against the labels you chose. Three sectors in the bank: council, hospital, warehouse.
Seven question formats, the way Beyond Dictionary serves them. Every question carries layered hints — a nudge, the reasoning, then a deeper connection — so a wrong answer opens a door instead of closing one.
The follow-on questions — the ones that arrive once everyone agrees labels matter and the argument moves to which words.
Two things: a reader can predict what is behind it before clicking, and the destination then keeps that promise. NN/g's four tests name the ways it fails: specific (does it name the real destination), sincere (does the page deliver what was promised), substantial (does it still mean something lifted out of its sentence), and succinct (is it no longer than clarity requires). The four pull against each other deliberately — and there is no maximum word count, because concision is subordinate to the other three rather than above them.
Because it predicts the outcome. Bailey and Wolfson found that a correct first click is followed by task success around 87% of the time, while a wrong first click drops that to roughly 46%. Going back costs one click, so what people are actually losing is confidence rather than time. The practical consequence is that first-click testing is unusually cheap evidence — one screen, one task, no build — and it isolates the labels from everything else in the experience.
Because accuracy for the institution is not the same as usefulness for the visitor. "Bursary", "Revenues" and "Development Management" are all correct names, and every one of them requires the reader to already know how the organisation is arranged. Task names — pay your fees, council tax, planning and building — require nothing. Keep the department names where they belong, on the pages and in correspondence, and let the menu name the errand. Hospital signage has run both for decades: the department name with the plain words beneath it.
The substantial test is the one it fails. People scan for links rather than reading the sentences around them, and a screen reader can list every link on a page — where a column of identical "Learn more" entries names nothing at all. The repair costs nothing: replace the empty phrase with the destination's own name, so "Learn more" becomes "How refunds work". The surrounding sentence almost always reads better afterwards, because it no longer has to carry the meaning the link was supposed to.
Convert the opinion into a testable claim. Rather than debating whether "Solutions" is clear, run a first-click test: three or four real tasks, the current label against two contents-derived alternatives, with people who do not work on the product. It takes an afternoon and produces behaviour rather than preference — which matters, because people routinely prefer a label they then fail to use. It also changes the politics: nobody has to be wrong in a meeting, because the test is what turns out to be right.
Yes, when the problem is granularity rather than ambiguity. Film genre stopped discriminating once "Drama" held tens of thousands of unlike titles, and Netflix's response was to go finer rather than simpler — roughly 76,897 micro-categories, tagged by people working from a 36-page manual. Compare that with a supermarket, which handles coconut milk belonging in three aisles by stocking it in several places rather than inventing new categories. Both are correct responses to different faults: one vocabulary was too coarse, the other genuinely ambiguous.
Because vagueness is safe internally and expensive only externally. A container word like "Solutions" or "More" never has to be argued about, accommodates every team's content, and fits the grid neatly. The cost lands somewhere nobody in the room is sitting — on a reader deciding whether it is worth a click. That asymmetry is why label decisions need evidence rather than consensus: the people in the meeting are the only ones who already know what is behind the word, which is precisely what disqualifies their judgement.
Every number quoted above, with where it comes from and why it is here.
Four things changed in how you read a menu — and one place to take them next.
Structures fail in patterns — and the same levers that help people can be turned against them. The ethics chapter.