Why knowing every word isn't enough
A few years ago I got curious about learning another language, so I downloaded Duolingo. I first picked up Russian, then tried Korean, then Japanese, then Chinese.
It made me realise how easy learning a language could be. Everything seemed straightforward. The exercises all seemed doable, even though the language was completely new to me.
I settled on Japanese. I really wanted to travel to Japan, and I was hoping to get fluent enough to have proper conversations while I was there.
Then I realised how slow my progress was.
One day I spent a few hours on it and finished a whole level. I looked at what I'd actually learned from that level, and it was very little vocabulary. So I went and read what other people thought, and everyone seems to agree: it doesn't give you much vocabulary, and the drills repeat a lot.
I also found it a bit stressful. The streaks, mainly. Miss a few days and the bird starts yelling at you. You get a notification at ten at night telling you you're about to lose your streak, right when you're trying to wind down for bed. When I wasn't being consistent, that got to me.
So I decided there had to be something else. I did a lot of digging online, YouTube and forums, looking for the fastest way to actually learn this.
Riding on momentum
Here's what I found.
There's the language. And there's the writing system: thousands of characters called kanji, which you can't avoid once you're past beginner level.
So I thought, why not do the hard part now, while I've still got momentum? I went looking for the fastest way through the characters, and I found Heisig. In that method you learn to recognise the characters before you learn much of the language itself, by inventing a small story for each one so it sticks. I bought his book and drilled the characters in Anki, a free flashcard app. Every evening, every morning, and a chunk of my weekends.
The plan was a sprint. Get the hardest, most tedious part out of the way fast, then move on to the part I actually wanted, which is understanding people and talking to them.
Characters on their own don't get you anything, and I knew that. That was the argument for doing them quickly instead of dragging them out over years. Get through the stage where they're useless, and come out the other side where they aren't.
Now, Heisig doesn't teach characters in the order you're likely to meet them. It teaches them in the order that makes them easiest to memorise, building simple shapes up into complicated ones. I knew that going in.
The bet was that I'd get through all two thousand or so in a few months, close that chapter for good, and come out able to recognise the characters and write them. Then the only thing left would be speaking, and connecting those characters to the actual language.
I knew it was risky and I wanted to do it anyway. The risk is in the ordering. Stop early and you're not holding a proportional slice of the language. You're holding a pile of characters you'll hardly ever meet. The method pays off when you finish it.
I didn't finish it. I stopped at six hundred. Somewhat useful, still not useful enough.
After that I was a bit burnt out, and I took a break from languages for a few months.
During those months I watched a lot of videos about travelling to China. And I noticed Chinese has an advantage over Japanese: if you've got a character, it matches its sound.
So I switched. I went looking for the best tools for learning characters all over again, because I still believed characters were the thing to do first. Plenty of the videos I came across said the same: characters are so essential you should start with them.
A second job
I found a paid course I thought was really cool at the time. It promised you'd learn the characters and their sounds together, and be able to say them from day one, using a method built on making little movies in your head.
I signed up, and at first I genuinely enjoyed it. I was learning the sounds, which I hadn't been doing with Heisig, where it was recognition and nothing else. It felt like a step up. They give you words and vocabulary as you go, and some sentences to read. I learned about three hundred characters, and I learned them well.
But it was all built on top of a flashcard system.
If you haven't used one, here's how they work. Every new thing you learn becomes a card, and the cards come back for revision on a schedule. Skip a few days and you open it to hundreds waiting for you. And the faster you add new cards, the more revision stacks up ahead of you, because you have to clear the backlog before you're allowed to add anything new.
So I'd feel like I was building momentum, then have to stop and do the revision. Going back over things I'd already learned was boring. It felt like high-school homework, which isn't a feeling I miss. And it's not exactly gripping material: words, characters, scenes, props, all of it stuff I'd seen before.
It's a fairly impressive technique, I have to say.
It's also a big time investment, and the more you learn the more you forget. I gave it whole weekends. Before work, after work. It turned into a second job and it stopped being fun.
To be fair to the method, it works, and some people do it brilliantly. If you're on a break from school, or you're not working, you can put the hours in and those hours genuinely pay off. With a full-time job, a life and other hobbies, it's just enormous.
There was another problem too, this one specific to the scenes. I found it a struggle to invent ones that were memorable enough, and you can't do many in a day before they all blur together. Maybe a few every couple of days. Memory is just tricky like that.
Methods built on invented scenes and spaced repetition, where cards come back at longer and longer intervals, deliberately limit how much you take in. That's the design. Your brain can only hold so many new things at once, so the method rations them out.
Immersion says the opposite. Spend as much time in the language as you can, and your learning scales with the hours you put in.
I was trying to follow both, and they pull against each other.
Every input-based tool has made the same choice, and they were right to. But it's the thing that moved me. Learning a language shouldn't be painful. If your method only works when you grind, you've set yourself up to fail.
That isn't how anybody learns a language
Children in Japan and China can already speak before they go anywhere near characters. They spend years understanding people and talking back before they learn to write down what they can already say.
I'd done it backwards. I'd taken the part that comes last and put it first, with no language underneath it to attach it to.
It works better the other way round, and not just out of tradition. Once you can follow the language and say things in it, characters stop being abstract shapes you're memorising for a future you can't picture. They turn into a way of reading things you'd already understand if somebody said them out loud. There's a point to them.
Characters are a long game. Put them first and it's very easy to get discouraged and quit, which is exactly what I did. Twice.
So I moved to reading and listening to things I could mostly understand.
Then I hit a different problem
Here's something that sounds impossible until it happens to you. You can know every single word in a sentence and still have no idea what the sentence means.
That was me, constantly.
Every reading tool I tried worked the same way. Tap a word, get its meaning. That little translation you get back is called a gloss, and glosses are genuinely useful. I used them all the time. But I'd sit there with every word translated in front of me and still not know what the sentence was saying.
Japanese made it obvious. It has small words called particles that mark the job each part of a sentence is doing. Two of them, written は and を, have dictionary meanings you can look up. But try to find where they go in the English. They dissolve into the English word order and get swallowed into a chunk of it. Until you've seen that happen, you keep hunting for a purpose that isn't there.
You don't need Japanese to feel this, though. It happens in English too.
"She ran into an old friend."
Look up those words one at a time, the way a learner would. You get a woman moving at speed 🏃 and colliding 💥 with an elderly friend 👴. Every one of those definitions is correct. Put them together and they describe something that didn't happen.
What actually happened is much quieter. She met a friend by chance 👋. "Ran into" is a single expression with nothing to do with running, and "old" is about how long she's known him 📅, not his age. You know all of that without even thinking. Someone learning English doesn't, and nothing in that list of separate meanings is going to tell them.
So a gloss only ever describes one word at a time. And that turns out to be wrong in two different ways at once.
The first is that words have jobs. In any sentence, the words are doing things in relation to each other, and a list of separate meanings tells you nothing about what those things are. That was the part I could never get to. I had all the meanings and still couldn't tell what the sentence was doing with them.
The second one is stranger, and it took me a while to accept. A lot of what looks like several words is actually one.
Linguists have been finding this since the 1980s, starting with Andrew Pawley and Frances Syder: around half of what comes out of our mouths isn't built word by word at all [6][7]. It arrives in ready-made blocks, stored and pulled out whole. "Ran into" is one of them.
So when an app splits a sentence into separate words, it's misrepresenting about half of it.
Two problems, then. Now the good part.
Someone already solved this. About a century ago.
Not in an app. In academic papers.
Linguists have this problem professionally. They write about languages their readers don't speak. So "here are the words, good luck" was never going to work for them either.
Their fix is over a hundred years old. It's called interlinear glossing, and you could draw it on a napkin.
You write the sentence. Then, lined up underneath, you put the meaning, so you can see which part goes with which.
Here it is in Spanish:
Se me olvidó el libro I forgot the book
Two Spanish words in that first column, one English word underneath, and the English word is "I". There is no "I" anywhere in the Spanish. It's carried by "se me".
The Spanish is quietly doing something else as well. Nobody gets blamed. The book forgot itself and I was merely nearby. That's the sentence a Spanish speaker actually reaches for.
Now here's what that same sentence looks like in a normal reading app. You tap "se" and you get its dictionary entry: himself, herself, itself, yourself, themselves, oneself. You tap "me" and you get "me". You tap "olvidó" and you get "forgot".
Line those up and you have: itself, me, forgot, the book.
Nobody in that sentence forgot anything to themselves, and there's still no "I" in sight. Every one of those definitions is correct, and together they take you nowhere. "Se" is one of the hardest-working words in Spanish, doing at least five different jobs depending on the sentence, and the dictionary can't tell you which one you're looking at right now.
Here's another, going the other way:
Mis amigos vinieron a mi casa a cenar My friends came over for dinner
Look at the middle column. "Came over" in English hides where they came to. Spanish says it out loud: they came to my house. Four words on top, two underneath, same meaning.
Spanish rearranges who's doing what. Japanese goes further. Whole runs of Japanese words collapse into a single English one, and the small words that hold the sentence together get folded in with them.
、。
Laid out flat, that sentence is a table. It's the table your head would otherwise have to build.
| The English | The Japanese under it | What that tells you |
|---|---|---|
| tap | タッチ + し + ます | Three Japanese words, one action. Not three meanings to assemble |
| the station, | 駅 + に | 駅 is the station. に marks it as the place you're going into, and English has no separate word for that, so it's folded in |
| your card | カード + を | Same again. を marks the card as the thing being tapped, and it's folded in too |
| at | で | The same kind of small word, and this one does come through as an English word of its own |
| When you | とき | One Japanese word opening out into two English ones |
| the ticket gate. | 改札 | Fifth word in the Japanese, last thing the English says |
Those small words are called particles, and they're the reason a dictionary can't get you through a Japanese sentence. Look up に on its own and you'll get a list of jobs it might be doing. Which one it's doing right here, in this sentence, is the thing you can see on the line and nowhere else.
That's the whole idea. That's the thing I'd been missing all along. Not more translation. Just being shown which piece goes with which, on the page, instead of rebuilding it in my head every time.
It's even standardised. The conventions most linguists use are called the Leipzig Glossing Rules [2], published jointly by the Max Planck Institute for Evolutionary Anthropology and the University of Leipzig. People who work with unfamiliar languages every day needed this badly enough to agree on a standard for it.
So why should being shown the structure help you learn?
Richard Schmidt, a linguist at the University of Hawai'i, put forward an idea that's now well known in language research. He called it the noticing hypothesis [3], and not everyone in the field agrees with it [4][5]. It says you only learn the parts of a language you consciously register while reading or listening. If your eyes pass over something without you registering it, it doesn't go in.
If that's right, then making the structure visible is the part that does the work.
That's exactly what happened to me. Once the structure was visible, I stopped needing rules explained, and the patterns started assembling themselves.
So I built it
That second line, the one under the Spanish, is the thing I wanted while I was reading. Not in a paper. On a screen, in a story, while I'm actually reading it.
Here's what that turned into.
Every word carries the colour of the meaning it belongs to, so the links are visible before you touch anything. Then touch a word, or hover it on a computer, and that word and its counterpart highlight together. Same idea as the lines above, except you don't have to hold any of it in your head. It's sitting there while you read.
The meaning line is a toggle. Turn it off and you're reading the language on its own, which is the point of all this. Turn it on the second you're stuck, and the interlinear glossing appears underneath, taking the sentence apart piece by piece. That switch is more useful than I expected: it turns every line into a test you can mark instantly.
There's a small margin of error in those links. Getting them exactly right across every language is hard. But what you're after is a grasp of where each word is actually going, from your own language into the one you're learning. Having that roughly right is worth far more than not having it, and not having it is what you get nearly everywhere else.
If you've seen coloured words in a reading app before, they were most likely telling you which words you already know. These colours aren't about you. They tell you what each word is doing in the sentence you're looking at.
Groups of words that only mean something together are marked as one piece, with what the whole thing means and when you'd use it. Like "ran into". That's the second problem answered. Splitting it into "ran" and "into" would be teaching you something false about the language.
Porfor / through / in (por la mañana)preposition lathe (fem.)article mañanamorningnoun abroto openverb lathe (fem.)article ventanawindownoun
Grammar patterns are marked where they show up, with a short explanation on the line and an example taken from the story you're reading. One per line, deliberately. Any more and you stop reading a story and start reading a textbook.
Porfor / through / in (por la mañana)preposition lathe (fem.)article mañanamorningnoun abroto openverb lathe (fem.)article ventanawindownoun
That last one isn't a research finding, it's a judgement call. Usually you can half tell what a bit of grammar is doing just from the sentence it's sitting in. You don't need teaching. You need somebody to confirm it, quickly, so you can carry on.
All of this sits on the line, rather than behind a search box, for one reason.
The moment you leave the story to look something up, you've stopped reading. You open a tab, you read an explanation built around somebody else's example sentence, you come back, and the thread is gone. Do that four times in a paragraph and you'll manage two lines before you quit and do something else. That isn't a discipline problem. That's what happens to everybody, and it's why so much language study feels tedious and academic and a bit old-school.
There's a second cost too, and it's the bigger one. Every minute you spend looking things up is a minute you're not spending in the language. Time in the language is the thing that actually builds it in your head. Interruptions eat that time.
And the other half of it is much simpler. Stories have a next bit. You want to know what happens, and that pull carries you through far more sentences than willpower ever will.
Which is where this comes back to the interlinear glossing.
On paper it only ever worked one sentence at a time, in a journal, for a linguist studying it. What's different now is that you can have it running underneath a whole story, live, while you read. No tabs. No looking anything up. No guessing at what a word is doing and hoping you got it right. You register the structure as you go, and you don't stop moving to do it.
Every script gets a reading layer over the top of it. If you can't read the writing system yet, none of the rest helps you, so there's a guide sitting above the text telling you how it sounds. Japanese gets furigana or romaji, Chinese gets pinyin, Arabic gets a transliteration. Different names per language, same job.
Anything you tap, you can keep. Words, phrases, grammar patterns. It all collects in one place you can look back at, and everything keeps its reading layer, so a word you saved three stories ago is still readable when you come back to it.
Porfor / through / in (por la mañana)preposition lathe (fem.)article mañanamorningnoun abroto openverb lathe (fem.)article ventanawindownoun
You can also keep the pairings themselves. Take a piece of the meaning line and save it together with the words it maps to, and what you've kept is a real bit of language from a real story with its context still attached. Not an entry off a list. Something you saw somebody actually say.
Porfor / through / in (por la mañana)preposition lathe (fem.)article mañanamorningnoun abroto openverb lathe (fem.)article ventanawindownoun
One more decision, and this one does have research behind it.
Everything is explained in your language, not in simplified versions of the language you're learning. Plenty of tools do the opposite, on the theory that you should stay immersed. So which works better?
Somebody checked. A 2024 review in the journal Language Teaching Research [1] pulled together the studies that had compared the two approaches head to head. Both help you learn. Explaining a new word in the reader's own language helped more. And it helped beginners most of all, which is exactly who this is for.
Why the stories are graded, and why words keep coming back
Reading without help only works once you already know most of the words on the page. In 2000 two researchers, Marcia Hu and Paul Nation, put the threshold at around 98% [8], which is higher than most people expect. It's the benchmark everyone uses, though a later replication couldn't fully reproduce it [9].
That explains something every beginner runs into. Authentic material isn't hard because you're slow. It's hard because too many words in every sentence are missing, and no amount of effort fills that gap.
So the stories are graded. They keep you above that line, and the line moves up as your vocabulary does.
Vocabulary itself doesn't arrive in one go either. Research on this has found that you need to meet a word several times, in context, before it starts to stick, and more times again before you really know it [10].
So words and phrases keep coming back to you. You'll meet the same word a few stories later, in a different sentence, doing a slightly different job. Sometimes you'll hear it before you spot it in the text. Each time you meet it again, it settles a bit deeper.
There's a library of stories in every language, and it keeps growing.
Listening while reading, then saying it out loud
The audio is synced to the transcript, so the line you're hearing lights up as it's read. You control the playback while you read along, and words and phrases light up individually when you touch or hover them.
Tambiénalso / tooadverb haythere is / there are (hay)verb unaa / an (fem.)article sillachairnoun yandconjunction unaa / an (fem.)article lámparalampnoun
There is also a chair and a lamp.
Porfor / through / in (por la mañana)preposition lathe (fem.)article mañanamorningnoun abroto openverb lathe (fem.)article ventanawindownoun
In the morning, I open the window.
You hear where the stress falls, where the voice rises and where it drops at the end of a sentence. In Mandarin you hear the tones properly, which is not a small detail when the tone is the difference between two entirely different words.
I didn't expect that to matter as much as it does. I remember more of what I read when I'm listening to it at the same time. It also makes shadowing straightforward, if you want to say the lines along with the recording as it plays.
After each story there are exercises, built out of what you've just read while it's still fresh. You use the story's patterns to make sentences about your own life. You translate, your language into the one you're learning. And then you give a short speech, about a minute, on an angle you pick, with an arc: open, explain, give an example, say what you think, contrast it with something, close.
That arc is the whole decision. Textbooks make you speak in drills: sentences nobody would say, about people who don't exist, in situations you'll rarely be in. Humans tell each other stories, and have done for as long as there have been humans. So the practice is shaped like something you'd tell a friend, not something you'd repeat back to a teacher.
I've started using it before sessions with my language tutor, which wasn't something I designed for. I put a short speech together about the story I've just read, and then I can actually talk to them about it, because the stories are about things worth talking about in the first place.
What gets me every time is what they notice. Not just that there are new words, but that the words fit. They turn up in the middle of a sentence where they belong, instead of being dropped in because I learned them last week. Textbooks have a strange habit of handing beginners vocabulary they'll almost never need, and you only find out how strange when you try to use it on a real person.
Where flashcards do belong
None of this means spaced repetition doesn't work. It works.
My problem was never the mechanism. It was building a whole method on top of it, where a revision queue decides how fast you're allowed to move.
As a supplement, it's a completely different thing. So everything you save exports to Anki, the free flashcard app, in a way that updates your existing cards instead of creating duplicates every time you export.
lingotide-japanese.apkg
4 cards
One click downloads a real .apkg — double-click to import into Anki Desktop, AnkiDroid or AnkiMobile. Base → target, the reading kept, the category as a tag; stable IDs mean a re-export updates your cards, never duplicates.
deck · LingoTide::Japanese
市場
いちば
ichiba
market
I deliberately didn't build a review queue into LingoTide. Anki is free, it's better than anything I would have written, and your cards should outlive my product.
Who it's for
If you're a beginner, or somewhere on the way to intermediate, and still building up the vocabulary to read anything real.
If you've got three apps, a textbook and a stack of flashcards, and none of them know about each other.
If drills bore you, and you keep losing momentum because of it.
And if what you actually want is to spend time inside the language rather than studying it from the outside.
That's who I built it for, because that was me.
Why I use it myself
I built this for me first, and I still use it every day. Here's the honest version of what that's like, and why I think it all comes back to one thing.
It's relaxing. That's the part I didn't plan for, and I think I know why it happens. I'm never stopping. There's no moment where I hit a sentence I can't unpick and go off looking for an explanation, because the interlinear glossing is already sitting underneath it. That second line, the one linguists have been writing under sentences for a hundred years, running live under every sentence I read.
That's the difference between this and the other reading tools I used, and it isn't a small one. Plenty of them will give you the words. What they don't give you is what the words are doing to each other, so you're still assembling that yourself, sentence after sentence, all the way through. That assembly is the part that wears you out. Take it away and reading in a language you barely know stops feeling like work.
And because nothing interrupts, the time actually counts. The more I put in, the more I remember. That wasn't true of anything else I tried, where more time mostly meant a bigger revision backlog waiting for me.
It's calm enough that I read before bed and fall asleep straight after, then pick it up again in the morning before anything else. On the train. Waiting for things. Anywhere I've got ten minutes.
I don't need anything else running alongside it either. What I save stays saved, it syncs out to Anki when I want flashcards, the exercises are there, and the next story is there.
There are free stories in every language, and they don't need an account: browse the free stories
If you try it, tell me where the links confused you. That's the part I'm still improving.
References
- [1]Kim, H. S., Lee, J. H. & Lee, H. (2024). The relative effects of L1 and L2 glosses on L2 learning: a meta-analysis. Language Teaching Research 28(1), 7-28. https://doi.org/10.1177/1362168820981394
- [2]Comrie, B., Haspelmath, M. & Bickel, B. The Leipzig Glossing Rules. Max Planck Institute for Evolutionary Anthropology and University of Leipzig, last revised 2015. https://www.eva.mpg.de/lingua/resources/glossing-rules.php
- [3]Schmidt, R. (1990). The role of consciousness in second language learning. Applied Linguistics 11(2), 129-158. https://doi.org/10.1093/applin/11.2.129
- [4]Schmidt's hypothesis is disputed. Tomlin, R. & Villa, V. (1994), Studies in Second Language Acquisition 16(2), 183-203, separate detection from awareness. https://www.cambridge.org/core/journals/studies-in-second-language-acquisition/article/abs/attention-in-cognitive-science-and-second-language-acquisition/16CF2D3A83B807507C10341287956D07
- [5]Truscott, J. (1998). Noticing in second language acquisition: a critical review. Second Language Research 14(2), 103-135. https://doi.org/10.1191/026765898674803209
- [6]Pawley, A. & Syder, F. (1983). Two puzzles for linguistic theory: nativelike selection and nativelike fluency. In Richards & Schmidt (eds), Language and Communication.
- [7]Erman, B. & Warren, B. (2000). The idiom principle and the open choice principle. Text 20(1), 29-62. https://doi.org/10.1515/text.1.2000.20.1.29
- [8]Hu, M. & Nation, P. (2000). Unknown vocabulary density and reading comprehension. Reading in a Foreign Language 13(1), 403-430. https://nflrc.hawaii.edu/rfl/item/43
- [9]Kremmel, B. et al. (2023). Unknown vocabulary density and reading comprehension: replicating Hu and Nation (2000). Language Learning. https://doi.org/10.1111/lang.12622
- [10]Webb, S. (2007). The effects of repetition on vocabulary knowledge. Applied Linguistics 28(1), 46-65. https://academic.oup.com/applij/article-abstract/28/1/46/174744