Learning a language by reading: what the research shows

You can know the meaning of every single word in a sentence and still not understand the sentence. That's the problem LingoTide is built around, and this page lists the research behind each piece of how it works, not just the claim.

Glosses in your own language, not the target language

A gloss is a word's meaning, shown right next to the word instead of behind a dictionary lookup. There's a substantial body of research on this, and the most useful piece of it is a 2024 meta-analysis by Kim, Lee and Lee that pooled 78 effect sizes across 26 studies [1]. Glosses written in a learner's own first language beat glosses written in the language being learned, with a medium effect size. Both helped. The first-language ones helped more, and the effect was largest for beginners.

That's why LingoTide explains every word in your own language rather than in simplified target-language definitions. It's the better-evidenced choice, not just the easier one to build.

One thing this research doesn't cover: every study behind it measured an ordinary gloss, a word next to its meaning. None of them measured a synchronised, colour-matched mapping running under a whole sentence, which is what LingoTide actually does. Nobody has studied that specific version yet.

Why a visible mapping matters at all

Knowing every word's meaning still doesn't tell you which word is doing which job in the sentence. Linguists have a name and a formal notation for showing that: interlinear glossing, standardised in 2003 by the Leipzig Glossing Rules, from the Max Planck Institute for Evolutionary Anthropology and the University of Leipzig [2]. It's a notation standard, not an efficacy study — it says scholars have formalised writing a sentence's structure directly underneath it for a long time; it says nothing about whether that helps a learner.

The closest thing to a theory for why a visible structure would help learning at all is Schmidt's noticing hypothesis: you acquire what you actually notice in the input, not everything you're merely exposed to [3]. It's one influential line of research in the field, not a settled fact.

LingoTide's version of this is the mapping under every sentence: each part of the target language is matched by colour to its meaning, and touching either end highlights both together. The colour marks what belongs together at rest; the highlight is what happens when you look closer.

Fixed phrases, not word by word

Roughly half of fluent native speech and writing isn't assembled word by word. It runs on stored multi-word units, pulled out whole, the way you'd reach for a ready-made package rather than building one from parts each time. Pawley and Syder made the argument in 1983 [4]; Erman and Warren put a number on it in 2000 and later estimates land around half, depending on how the counting is done [5].

Treating a sentence as a bag of individual words is wrong about half the time by this measure. So LingoTide marks a fixed expression as one package, not two or three separate words, with what it actually means and when you'd use it.

Reading at a level you can actually follow

The standard benchmark for how much of a text you need to already know is Hu and Nation's 2000 study: roughly 98% known-word coverage for comfortable, unassisted reading, which works out to several thousand word families for unsimplified writing [6]. The figure is more fragile than its popularity suggests — it came from a small sample reading a short text with invented words swapped in, and a 2023 replication couldn't fully reproduce it [7]. It's still the most widely used number in the field, because it names exactly what a beginner runs into with authentic content: it isn't hard because you're slow, it's hard because you're missing one word in five and no amount of focus fixes that on its own.

That's why LingoTide's stories are graded rather than authentic. Grading keeps you above that line while your vocabulary grows, instead of throwing you at native-speed content before you're ready for it.

Meeting a word again, instead of drilling it

Webb ran a 2007 study where learners met unknown words at 1, 3, 7 or 10 encounters in context [8]. Learning was already considerable by ten encounters, and Webb's own conclusion was that full knowledge of a word takes more exposure than that. Repetition in context works; ten repetitions is where the gains become worth having, not where they stop.

That's the reasoning behind the corpus design: the same words resurface across different stories, in different sentences, doing a slightly different job each time, instead of sitting in a queue you have to schedule yourself through. When you do want spaced repetition on top of that, everything you save exports to a real Anki deck with stable card IDs, so exporting again updates your existing cards instead of duplicating them.

Why the app asks you to speak

After every story there's a prompt: say something out loud in the language you're learning, then check it against a model answer. The reasoning behind that comes from Swain's 1985 study of French immersion students in Canada, who had years of strong comprehension and still made systematic errors in production [9]. Her argument was that understanding a message and producing one require different things, and that being pushed to produce something precisely, not just well enough to get by, is what closes that gap.

It's a hypothesis built on classroom observation, not a controlled trial of speaking prompts specifically, and nobody has tested LingoTide's particular version of it. But it matches something almost every learner notices about themselves: you can understand far more than you can say.


There are free stories in every language, and they don't need an account: browse the free stories

If you try it, tell me where the mapping confused you. That's the part still being improved.


References

  1. [1]Kim, H. S., Lee, J. H. & Lee, H. (2024). The relative effects of L1 and L2 glosses on L2 learning: a meta-analysis. Language Teaching Research 28(1), 7-28. https://doi.org/10.1177/1362168820981394
  2. [2]Comrie, B., Haspelmath, M. & Bickel, B. The Leipzig Glossing Rules. Max Planck Institute for Evolutionary Anthropology and University of Leipzig, last revised 2015. https://www.eva.mpg.de/lingua/resources/glossing-rules.php
  3. [3]Schmidt, R. (1990). The role of consciousness in second language learning. Applied Linguistics 11(2), 129-158. https://doi.org/10.1093/applin/11.2.129
  4. [4]Pawley, A. & Syder, F. (1983). Two puzzles for linguistic theory: nativelike selection and nativelike fluency. In Richards & Schmidt (eds), Language and Communication.
  5. [5]Erman, B. & Warren, B. (2000). The idiom principle and the open choice principle. Text 20(1), 29-62. https://doi.org/10.1515/text.1.2000.20.1.29
  6. [6]Hu, M. & Nation, P. (2000). Unknown vocabulary density and reading comprehension. Reading in a Foreign Language 13(1), 403-430. https://nflrc.hawaii.edu/rfl/item/43
  7. [7]Kremmel, B. et al. (2023). Unknown vocabulary density and reading comprehension: replicating Hu and Nation (2000). Language Learning. https://doi.org/10.1111/lang.12622
  8. [8]Webb, S. (2007). The effects of repetition on vocabulary knowledge. Applied Linguistics 28(1), 46-65. https://academic.oup.com/applij/article-abstract/28/1/46/174744
  9. [9]Swain, M. (1985). Communicative competence: some roles of comprehensible input and comprehensible output in its development. In S. Gass & C. Madden (Eds.), Input in Second Language Acquisition (pp. 235-253). Newbury House.
Learning a language by reading: the research | LingoTide