Every product is a building people walk through without a guide. Information architecture is the decision about which rooms exist, what is written on each door, and which room sits inside which. Get it right and nobody notices. Get it wrong and no menu, however handsome, will rescue it.
Describe an app you use daily without mentioning a single colour. Most people end up listing things — orders, account, settings, help — and then saying which sits inside which. That list and that nesting are the information architecture: what exists, what each thing is called, what contains what. Every one of those decisions is settled before a menu is drawn, and each one outlives the menu’s next redesign. Navigation is the visible control on top; architecture is the building underneath. Mistaking one for the other is why so many teams answer “nobody can find anything” with a new menu bar, then find that nothing improved.
People do not read a site, they forage through it — scanning for a word that smells like the thing they came for, clicking it, then checking whether the next screen still smells right. Researchers call this information scent, and it accounts for most of the difference between a structure that works and one that does not. A strong label lets someone predict what is behind it before clicking, so the click costs nothing. A vague one makes every click a gamble, and a few lost gambles send people to the search box or out of the door. Labels are not decoration applied to the structure. To the reader, they are the structure.
Five top-level labels on each side, with the same twelve pages sitting underneath both. Nothing was added or removed — the only change is what the groups are called and what went into each.
Paying a fee requires knowing in advance that fees are a Bursary matter. The visitor is being asked to supply knowledge only an employee has.
The same twelve pages sit underneath both columns. No inside knowledge is needed now — each label names a job somebody actually arrived to do.
Joshua Porter tested this properly for UIE in 2003: 44 people, 620 tasks, over 8,000 clicks.1 No drop-off in success after the third click. No fall in satisfaction. No correlation at all between how often someone clicked and whether they found the thing. Several participants visited as many as 25 pages and finished perfectly happy. What predicted success was confidence — whether each label confirmed they were still on the right trail. The rule survives because it is short and easy to measure, and it does real harm: teams flatten a working structure into one enormous menu to save clicks that nobody was counting.
George Miller’s 1956 paper on the span of short-term memory is where the number comes from — a measure of how many items anyone can hold in their head with nothing in front of them. A menu is in front of you. Reading it is recognition rather than recall, so the limit does not transfer. Miller later wrote that he had been “persecuted by an integer”2 and that calling the number magical was a rhetorical flourish, not a finding. Seventy years on it is still quoted to justify deleting useful options. Size a menu by how fast it scans and how precisely each label discriminates — never by a number borrowed from a memory experiment.
What follows is a real research method, not a simulation. A tree test strips away every visual — no colour, no layout, no logo — and asks one thing: given only the words, can a person find what they came for? Running it costs nothing because nothing has been built yet, which is exactly the point.
A useful working target is roughly 70–80% direct or eventual success3 — but the number moves with task difficulty, how mixed your audience is, and how critical the task happens to be. Compare mainly against your own previous version, and spend your attention on where the first clicks go wrong. That is where the structure is actually leaking.
Same four tasks, same site, same content. Only the top-level labels change. Pick where you would look.
The content never moved. Both menus lead to exactly the same pages — the only variable is whether the word on the door tells you what is behind it.
Flattening a structure does not remove work. It moves the work from stepping to scanning. Here is the trade, with 64 items to organise.
Simplified comparison model — not an empirical usability formula. It does not predict how long anyone takes; it only lets you compare two shapes of the same catalogue. What it shows matches practice: both extremes are expensive and the cheapest structures sit in the middle. In a real product label quality swamps this curve entirely — eight precise words beat four vague ones every time.
GOV.UK’s navigation was not designed and shipped. Six rounds of usability testing built it,5 starting from a prototype exposing only the lowest levels of the taxonomy and growing round by round. What came out: a breadcrumb on every page, topics in a grid, an accordion, and step-by-step journeys now reused across departments.
Baymard’s 2025 benchmark scored over 16,000 measurements across 180+ leading commerce sites. The most frequent navigation failure of all: 95% do not highlight where the visitor currently is.4 The cheapest thing a menu can offer is orientation, and almost everybody skips it.
A department reorganises, the category everyone bookmarked is renamed, and no redirects are written. Every saved link, every citation, every search result now lands on a 404. The pages still exist and the content is unchanged — but from outside, the section has vanished. An address becomes infrastructure the moment somebody links to it.
A tree tells you where pages live, but not what kind of page each node becomes. Without a small vocabulary of page types, teams preserve the sitemap and then reinvent the experience at every destination.
Between the sitemap and the screen sits a layer most teams never name: the set of page types. Once you can say that every page in a product is one of eight kinds, an enormous amount of argument disappears. Nobody invents a layout per page; they pick a type and fill it. The structure becomes something a team can hold in its head, and a new page arrives already knowing what blocks it needs, how it behaves on a phone, and what it owes the reader. The rule that makes it work is short: no page ships without matching a type.
Everything so far has been one principle at a time. This is all of them at once, on a menu that is genuinely broken.
Seven question formats, the way Beyond Dictionary serves them. Every question carries layered hints — a nudge, the reasoning, then a deeper connection — so a wrong answer opens a door instead of closing one.
The follow-on questions — the ones that come up once the principle is agreed and the argument moves to what to actually do.
Redesigning the navigation is right when the structure tests well and people still get lost. That happens more often than you would think: a sound tree can be hidden behind a menu that only opens on hover, buries a level, or never shows where you currently are. The test is simple. Run a tree test on the structure alone, with no interface. If people find things in the text-only tree but fail on the live site, the fault is genuinely in the navigation and a redesign will help. If they fail in the tree too, no menu will save it — you are looking at an architecture problem wearing an interface costume.
Ask whether each step confirms the reader is still on the right trail. That is the thing Porter's data actually pointed at, and unlike a click budget it is testable: run a tree test and look at where first clicks go wrong. A structure where people step confidently through five levels beats one where they hesitate at level two. The practical form is a question rather than a number — after this click, does the next screen tell the reader they guessed right? If yes, you can afford another level. If no, one more click is already one too many.
No number governs it, which is the honest answer and exactly why 7±2 filled the vacuum. What does govern it is how fast the list scans and how cleanly each label separates from its neighbours. Twelve precise, well-separated labels are easier than five that overlap, because overlapping labels force the reader to hold two candidates in mind and compare them — which is the expensive part. Judge a menu by the discrimination between its items, not by their count, and let a tree test tell you where two labels are competing for the same clicks.
Card sorting is generative and tree testing is evaluative. In a card sort you hand people your content and watch how they group it, which surfaces their mental model before you have committed to a structure. In a tree test you give them a text-only version of a structure you have drafted and ask them to find specific things, which measures whether that structure works. Sort first, draft from what you learned, then test the draft. Practitioners treat 70–80% task success as good, and often design toward 75% or better.
Probably the opposite. Heavy search use is as consistent with "browsing is broken here" as with "search is excellent here", and the first is more common. Search only works for someone who already knows the right word; browsing lets people recognise a label rather than recall a term. Cutting the menu removes the path that was serving everyone without the vocabulary, and it makes the metric look better while the underlying problem gets worse. Find out why browsing is being abandoned before you remove the alternative.
Not showing people where they are. Baymard's 2025 benchmark scored more than 16,000 UX measurements across 180+ leading e-commerce sites and found that 95% fail to highlight the user's current scope in the main navigation — the single most frequently cited failure in the whole set. The same benchmark found 58% of desktop sites and 67% of mobile sites scoring mediocre or poor on navigation overall. Orientation is the cheapest thing a menu can offer and the most commonly skipped, because when it is done well nobody notices it.
Because it looks user-centred and it matches how organisations picture their public. A menu that opens with "I am a… Student · Parent · Teacher · Employer" feels considerate. In use it adds a step before anyone can act, breaks for the many people who belong to two of the groups or none, and forces content that matters to several audiences to be duplicated and then maintained in parallel. People arrive with a task, not an identity. Structuring by what they came to do avoids the toll gate entirely.
Every number quoted above, with where it comes from and why it is here. An article that tells you to distrust unsourced rules of thumb owes you this.
Four things changed in how you look at a product — and one place to take them next.
You can read a structure. Now decide which device it has to win on — and why “mobile-first” is the wrong question to start with.