Cognates: The Vocabulary You Already Own

Chart showing the share of English words with Spanish cognates rising from 27% in the first thousand most frequent words to 85% beyond the third thousand.

The longest word in the paragraph is usually the one you already know. It's the five-letter ones that stop you.

You've been taught roughly the opposite — that a language gets harder as the words get longer, and that anything resembling English is a trap laid to catch you. Both halves are wrong. The second half is wrong by about three orders of magnitude.

Marta opens a Spanish article and her eye snags on inevitablemente. Fifteen letters. She looks it up. Two lines later she meets sacar — five letters, one of the most common verbs in the language — and has no idea. She closes the tab believing Spanish is hard. She has the map upside down, and so does almost every beginner course she will ever be handed.

Try the method first: Generate your first bilingual story

The Half of the Page Nobody Charged You For

Thomas Finkenstaedt and Dieter Wolff ran a computerised survey of roughly 80,000 entries in the Shorter Oxford Dictionary. French accounted for 28.3 per cent. Latin accounted for another 28.24. Germanic — the stock English is supposedly made of — came to 25.

More than half of the English dictionary is, by ancestry, Romance.

Dictionary counts flatter, though, because most of what's in a dictionary never appears in anything you read. The applied linguistics is narrower and more useful. In 2025, Thi Ngoc Yen Dang, Marijana Macis and Mireya Aguilera-Munizaga took the Academic Spoken Word List — 4,761 words distilled from thirteen million words of lectures, seminars and lab sessions across twenty-eight subject areas — and checked every item against Spanish. Half of them, 50.77 per cent, were cognates.

Sabrina Lubliner and Elfrieda Hiebert had already done the written equivalent and found 74.74 per cent of Coxhead's Academic Word List. Deborah Reed and colleagues went into biology textbooks and found 93.

Now turn it around. A cognate pair works in both directions. When you sit down with a Spanish text, something close to half the abstract vocabulary in it was installed years ago, free, as a side effect of learning English.

So why does none of it feel free?

Why the Easy Words Are the Hard Ones

Dang's team sorted their cognates by how common the English word is in ordinary use. The pattern is a staircase running the wrong way.

Among the thousand most frequent words of English, 27.08 per cent have a Spanish cognate. In the second thousand, 56.42. In the third, 78.64. Past the third thousand — the rare, long, technical-looking band — 85.44.

The more intimidating a word looks, the more likely you already own it. The everyday core is where the old Germanic stock survived: get, keep, shut, wet, run. That band has almost no free vocabulary in it at all.

Which explains something that otherwise reads as a personal failing. Beginner materials are built entirely from the high-frequency core — that is what makes them beginner materials. A graded reader capped at 400 headwords, a children's story about a dog lost in the mud: both are assembled from precisely the layer where your English is worth nothing. You were handed the one part of the language you get no discount on, and told it was the easy part.

There's a complication worth stating rather than hiding. Most of the cognates Dang's team found — 61.85 per cent — are common in English but relatively uncommon in Spanish. Reading Spanish, you'll meet them less often per page than the headline number suggests, and mostly in adult prose: essays, journalism, argument, anything making a case. Which is only the same point again from a different angle. The free vocabulary isn't waiting on the beginner shelf. It's one level up, in the material you assumed you weren't ready for.

The False-Friend Tax

Every cognate list ever published comes stapled to a warning about embarazada. Fair enough. Announce to a room that you're pregnant when you meant embarrassed and you'll be telling that story for a decade.

Now price the risk. In the Academic Spoken Word List, false cognates came to 0.06 per cent of the words — three items out of 4,761. Across the full thirteen-million-word corpus they covered 0.002 per cent of everything anyone said. Martínez found 2.33 per cent in engineering texts; Reed found 1 per cent in biology. Three decades of counting, in different languages of instruction and different genres, keep returning the same answer: false friends are a rounding error.

They may not even be a net cost. Marta Marecka and six colleagues, writing in Cognition in 2021, taught seventy-six Polish speakers twenty-four invented words paired with pictures. Some of the invented words resembled the Polish word for the object shown. Some resembled a different Polish word entirely — engineered false friends. Some resembled nothing. Real cognates were learned fastest, as expected. The false ones were learned as quickly as the unrelated words when participants had to recognise them, and faster when they had to produce them. The overlapping form helped even while the meaning it pointed at was wrong.

So consider what the standard warning actually asks of you. Distrust half the page, to insure against tripping over a fraction of one per cent of it. John Holmes and Rosinda Ramos named the failure mode in 1993 — reckless guessers — but recklessness is not the error most learners make. Refusal is. A wrong guess costs you one word. Suspicion costs you every word you were too careful to claim.

You Can Own a Word and Not Know It

In 1993 William Nagy, Georgia García, Aydin Durgunoğlu and Barbara Hancin-Bhatt tested seventy-four upper-elementary bilingual students who could read in both Spanish and English. The students read four expository texts seeded with cognate pairs — transform and transformar — and answered comprehension questions. Only afterwards was the concept of a cognate explained to them, and only then were they asked to go back through the texts and mark the words with Spanish relatives.

Two findings came out. Comprehension tracked cognate recognition: the students who spotted more understood more of what they'd read. And a substantial number of students failed to flag words they had already demonstrated, on separate vocabulary tests, that they knew in both languages.

They owned the connection. They weren't using it.

That is the whole problem compressed into a sentence. The vocabulary is inherited. Noticing it is a skill, and the skill does not arrive attached to the inheritance.

There's a catch about which sense it arrives through, too. Dang's team measured how much of each cognate pair survives in spelling versus in sound. Shared letters averaged 0.76 — three-quarters of the word. Shared phonemes averaged 0.52. Física and physics are the same word on paper and near-strangers out loud; nobody hearing física at conversational speed catches it. Your inheritance is a reading inheritance first. You'll see the family resemblance months before you can hear it.

What If Your Language Isn't Romance

Finland runs the cleanest natural experiment anyone has. Finnish-speaking and Swedish-speaking children, one country, one school system, the same English curriculum, the same age. Swedish is Germanic. Finnish is not related to English in any way at all.

Rainio tested both groups on the same English words. The Swedish speakers averaged 68 per cent. The Finnish speakers averaged 42. Palmberg had already identified the mechanism: of the twenty English words Swedish-speaking children understood most readily, fifteen closely resembled their Swedish equivalents. Sister and syster, solved by every child tested. Hand and hand, 95 per cent. Cat and katt, 89.

Håkan Ringbom, who spent a career documenting this, put it without decoration: if the target language is close to your own, you already possess a considerable potential vocabulary in it, and receptive command of a basic vocabulary can be reached with very little effort.

When the languages aren't close, the inheritance shrinks. It rarely disappears. Frank Daulton went through the Academic Word List looking for Japanese counterparts and found around 27 per cent — not shared ancestry but a century of borrowing, thousands of English words resident in Japanese and written in katakana. Japanese and English are about as unrelated as two languages get, and a quarter of the academic list is still sitting there pre-installed.

The question was never whether you have free vocabulary. It's how much, and whether you're collecting it.

Nobody can teach you the words you already own. You have to catch yourself knowing them, and that only happens mid-sentence, in text moving fast enough that considerable arrives as considerable and you don't stop. Bilingual stories put the translation in the next column, so the words that genuinely are new cost you a glance rather than a tab — and the ones you were about to look up out of habit reveal themselves as words you'd have understood anyway. Generate your first bilingual story at your level and pay attention to which words actually stop you. It won't be the long ones.


Related reading:

Start reading bilingual stories for free

Start reading bilingual stories for free

Not sure of your reading level? Take the free CEFR test →

Continue your reading practice

English Reading Practice A2 → Learn English Through Reading → English Short Stories for Beginners →

More from the blog

Extensive vs. Intensive Reading: Why Volume Beats Analysis
Why Children's Books Won't Teach You a Language
Fiction vs. News: What to Read When You're Learning a Language