Word Families: Why One Word You Know Is Really Twenty
You know the word use. The frequency list that told you so also counts usable, unused, misuse, user and about two dozen other forms as the same entry — one word, learned, ticked off.
That is either the best arithmetic in language learning or the reason your vocabulary number is fiction. Which one it is depends on something nobody ever tested you on.
Nagy and Anderson counted roughly 88,500 distinct words in the English that school children read. Nobody learns 88,500 separate things. What gets learned is a much smaller set of roots, plus a small set of parts that clip onto the front and the back of them — and the parts are where almost all the arithmetic happens.
Try the method first: Generate your first bilingual story.
What Counts as One Word?
A word family is a base word plus every form built from it by regular inflection and derivation. Care, cares, cared, caring, careful, carefully, careless, carelessness, carer — nine forms, one line on a frequency list.
Laurie Bauer and Paul Nation sorted English affixes into seven graded levels in 1993, using four questions: how often does the affix appear, how regular are its form and meaning, is it still being used to make new words, and can you predict what it does. Level 3 holds the affixes that pass all four cleanly: -able, -er, -ish, -less, -ly, -ness, -th, -y, non-, un-.
Ten parts. You have met every one of them a thousand times.
Nearly every word list, vocabulary test and "you need 8,000 words to read a novel" figure is built on this unit. The assumption underneath is that if you know care, then carelessness costs you close to nothing. Which is a claim about you, not about English.
A Handful of Affixes Covers Almost Everything
White, Sowell and Yanagihara went through the prefixed words in The American Heritage Word Frequency Book and found a distribution that is wildly lopsided. Un- on its own accounts for 26% of prefixed words. The top three — un-, re-, and in- meaning "not" — cover 51%. Add dis- and you reach 58%. Twenty prefixes get you to roughly 97%.
Suffixes behave the same way. Ten suffixes covered 85% of their sample, and three inflections — -s/-es, -ed, -ing — accounted for 65% of all the words carrying an ending.
Batia Laufer and Tom Cobb found the same shape in adult writing. In a corpus of roughly 250,000 words spanning academic texts, journalism, novels and graded readers, 22 affixes made up 98% of every affix that appeared.
Here is what that looks like in one sentence. Yulia is reading an English article about a city council and hits unenforceable. She has never seen the word. She knows force. Take off un-, take off -able, and what stands in the middle is enforce — and the sentence has already supplied the rest: a rule nobody can make people follow. No dictionary, no interruption, about a second and a half.
Nagy and Anderson put a number on this multiplier: for every word a learner knows, an average of one to three related words should also be understandable, the exact figure depending on how well that learner uses context and morphology.
Should is carrying a great deal of weight in that sentence.
Why the Deal Doesn't Close by Itself
Norbert Schmitt and Cheryl Boyd Zimmerman asked 106 non-native university students to produce the noun, the verb, the adjective and the adverb for sixteen words they had already demonstrated they knew. Producing all four was uncommon. Producing none was also uncommon. The typical result was two or three — a partly furnished family, with a couple of rooms nobody had been into.
You can watch this happen in your own reading. Decide is easy. Decision is easy. Then a sentence calls something a decisive week and you stall for half a beat on a word assembled entirely out of parts you already own.
Schmitt and Zimmerman were testing production, though, which is always harder than recognition — and reading only asks you to recognise. Stuart McLean removed that excuse. He tested 279 Japanese university learners on base forms, inflected forms and derived forms in a comprehension task, and the derived forms still came apart: significantly harder at every proficiency level, including for learners who had mastered the most frequent 1,000 words of English. His conclusion was blunt. Counting in word families inflates what a vocabulary score actually means, and research should switch to a smaller unit.
The family is real. It exists in the language. What does not automatically exist is your access to it — and the part of the deal you were never offered is the affixes themselves.
Where Affix Knowledge Actually Comes From
Masamichi Mochizuki and Kazumi Aizawa tested 403 Japanese learners on thirteen prefixes and sixteen suffixes, attached to invented words so that nobody could answer from memory of a specific vocabulary item. The average learner had a vocabulary of about 3,769 words and knew 7.24 of the prefixes and 10.70 of the suffixes.
The distribution is the interesting part. Learners in the 2,000-word band understood 45% of the affixes tested. In the 3,000-word band, 61%. In the 4,000-word band, 70%. Affix knowledge climbed in step with vocabulary size, correlating at 0.65.
That is a correlation measured at a single point in time, so it does not prove that reading installs affixes; larger vocabularies and better affix knowledge could both be products of more study, or simply of more of everything. What it does rule out is the comfortable idea that morphology is a separate module you bolt on later. It arrives alongside the words, roughly in proportion to how many you have met.
And it arrives the way patterns always do. You do not learn -ness from a definition of -ness. You meet darkness in one story, kindness in another, thickness in a third, and at some point the ending stops being noise attached to a word and starts being a thing that does a job.
The Texts That Starve You of Derived Words
Laufer and Cobb also measured how much of a text is made of derived forms. Newspaper articles: 7.88% of running words. Academic texts: 7.78%. Classic novels: 5.04%. Graded readers: 3.17%.
Per thousand words, that is about 79 derived forms in journalism and about 32 in a graded reader. Spend a year inside simplified text and you will have seen the machinery less than half as often as someone reading the real thing, while believing you were doing the careful, level-appropriate version of the same activity.
Laufer and Cobb drew a reassuring conclusion from their own numbers: you do not need most of a word family in order to read, because base words, inflections and a small set of frequent affixes will carry you past the coverage you need. Other researchers have pushed back hard on that, arguing derived forms are neither so rare nor so transparent for anyone who is not already advanced. The argument is unresolved. Nobody in it claims that graded readers give you more exposure to affixes than novels do.
Which points to a way of reading rather than a thing to study. When a long word stops you, go for the root before you go for the dictionary: take off the front, take off the back, see what is standing in the middle. If it is a word you know, keep moving. If it is not — if you strip invite down to vite, which is nothing — then you have found a genuine gap, and that is the word worth the interruption.
The vocabulary you are building is not a list. It is a modest set of roots wearing a much smaller set of coats, and the coats repeat so relentlessly that a few dozen of them account for nearly everything you will ever read. That is not knowledge you can memorise into place, because it is not facts — it is pattern recognition, and pattern recognition runs on volume. Generate your first bilingual story at your level and see how the method feels. Read enough real sentences and one day unrepeatable will not be a word you look up. It will be three things you already know, standing next to each other.
Related reading:
- Cognates: The Vocabulary You Already Own
- The Pickup Rate: How Few Words a Book Actually Teaches You
- Vocabulary Size: How Many Words You Actually Need to Read
Start reading bilingual stories for free
Start reading bilingual stories for freeNot sure of your reading level? Take the free CEFR test →