Spaced repetition is the least controversial idea in language learning and one of the most poorly implemented. The principle is simple enough to state in a sentence, and almost every part of applying it well is counterintuitive.
This explains what it is, why the obvious way to use it is wrong, and what a good implementation inside an AI language app looks like.
The principle
Memory decays predictably. Review an item just before you would have forgotten it, and the memory strengthens far more than if you had reviewed it while it was still fresh. Each successful review extends the interval — a day, then three, then a week, then a month.
The corollary is the part people resist: reviewing something you remember well is close to worthless. The effort that produces retention is the effort of nearly failing to recall. Comfortable review feels productive and does very little.
This is why spaced repetition is a scheduling problem rather than a content problem, and why a computer does it better than a person. Tracking thousands of items, each with its own decay curve, is exactly the kind of bookkeeping no human tutor could manage.
Where flashcard-style implementations fall short
Most apps implement spaced repetition as a card: the word on one side, the translation on the other. It works, within limits, and the limits are significant.
It trains recognition, not production. Seeing hablar and recalling "to speak" is a recognition task. Producing hablar when you need it mid-sentence is a different operation, and passing the card does not demonstrate you can do it.
It strips context. Words do not behave identically in every setting. A card teaches a word-to-word mapping and leaves out which prepositions follow it, which register it belongs to, and which words it habitually appears beside — most of what "knowing a word" actually means.
It measures the wrong success. The card is marked correct when you recall the translation. But translation is not the goal; deployment is. Many learners have thousands of mature cards and still hesitate mid-sentence, which is not a failure of the algorithm — it is a mismatch between what was tested and what was wanted.
What a better implementation looks like
Self-assessment rather than automatic marking. You know whether a word arrived instantly or was dragged up with effort, and that distinction matters more than right-versus-wrong. A system that asks you to judge recall quality gets better scheduling data than one inferring it from a binary answer.
Explicit progress rather than opaque state. Knowing a word sits at 30% mastery and will return in a few days is more motivating than a hidden algorithm, and it lets you see the difference between words that are genuinely settling and words you keep re-failing.
A defined finish line. Items should retire. Without a mastery threshold, review load grows without bound and the deck becomes something you maintain rather than something you learn from.
Connection to production. The best case is when scheduled vocabulary shows up in conversation, because using a word in a sentence you constructed is a far stronger memory event than recalling it from a card.
How Enverson AI implements it
Enverson AI's Vocabulary tab is built on the self-assessment model. You begin by sorting which words you already know, so the schedule starts from your actual position rather than a generic beginner baseline.
Review uses a swipe: indicate whether you remember a word or not. Words you do not remember return sooner. Words you do remember advance — a step at a time, for example to 30% — and return after a few days. Each word climbs toward 100%, at which point it is marked learned and stops consuming review time.
Two things about that design are worth noting. The visible percentage turns an invisible algorithm into something you can reason about. And the mastery threshold means the deck has a finish line, so review load does not accumulate indefinitely.
What connects it to production is the rest of the app. Enverson AI's Multidimensional Personalization Engine (MPE) tracks vocabulary range as one of six independent dimensions — alongside grammatical accuracy, speaking pace, fluency, filler-word frequency and conversational complexity — and a Free Talk session reports which words you actually used, not which you could recognise.
That closes the recognition-production loop that flashcard apps leave open. A learner can see their vocabulary mastery climbing in the Vocabulary tab while their vocabulary range in Free Talk stays flat, and that gap is precisely the diagnosis: words are being learned in a form that is not reaching speech.
Its limits: learning is mobile-only (iOS and Android), and it supports English, Spanish, German, French and Russian.
Duolingo handles review well within its course structure, and Babbel pairs vocabulary with clear grammatical explanation — but neither reports whether a learned word is appearing in your unscripted speech.
The three intervals that matter most
Most of the benefit comes from a small number of early reviews, which is worth knowing because it changes how you allocate attention.
The first review, within a day. Forgetting is steepest immediately after learning. An item reviewed once within twenty-four hours survives dramatically better than one first seen a week later, and this single review does more work than any subsequent one.
The three-to-four day review. The point at which most learners would begin to lose an item. Catching it here is what converts a word from short-term familiarity into something durable.
The two-week review. Roughly where a word either settles permanently or reveals that it never properly formed. Items that fail here usually fail because they were learned as isolated translations rather than in usable context.
After that the intervals stretch and the marginal value of each review drops sharply. A word at 80% mastery does not need your attention nearly as much as a new one does — which is why systems that keep old items circulating indefinitely waste time that new material would repay better.
Why "I keep forgetting this word" happens
Almost every learner has a handful of words that refuse to stick despite dozens of reviews. The cause is usually one of three things, and none of them is repetition count.
No hook. The word has no connection to anything you already know — no cognate, no memorable sound, no context in which you first met it. Isolated facts decay fastest. The fix is to attach it to something: use it in a sentence about your own life, or learn it alongside a word it habitually appears with.
Interference. Two similar words learned close together compete, and each review of one weakens the other. Learning them apart, or deliberately practising the contrast, resolves it faster than more repetitions of either.
It was never learned in a usable form. If you only ever met the word as a translation pair, you have stored a mapping rather than a meaning. Encountering it in speech — hearing it, then producing it in a sentence you built — creates a different and much more durable trace.
How to use spaced repetition well
Be honest when self-assessing. Marking a word remembered because you nearly got it corrupts the schedule and costs you the word later. The uncomfortable judgement is the useful one.
Keep sessions short and daily. Ten minutes every day outperforms an hour on Sunday, because the algorithm assumes you appear when items are due. Skipping days means reviewing things you have already forgotten, which is inefficient rather than harmless.
Learn words you will use. Frequency beats interest. A thousand common words carry ordinary conversation; a thousand words chosen because they were interesting will not.
Do not let it become the whole practice. Vocabulary review is support work. It is comfortable, quantifiable and satisfying, which makes it a very effective way to avoid speaking — and speaking is where the words become usable.
Learning words in the wrong unit
A structural point that undermines a great deal of vocabulary study: the word is often the wrong unit to learn.
Fluent speech is substantially built from multi-word chunks that behave as single items — fixed collocations, common phrases, verb-plus-preposition pairs. A speaker does not assemble these from parts each time; they retrieve them whole. Learning the components separately and hoping they combine correctly produces sentences that are grammatical and subtly wrong, because the combinations are conventional rather than derivable.
The practical consequence is that a card pairing one word with one translation teaches something narrower than intended. Make and do both translate identically into many languages, and knowing both individually does not tell you which one goes with a decision.
Better practice is to learn the chunk. Not the verb alone but the verb with its habitual partner; not the noun alone but the noun with the article and the preposition that usually follows. It feels like more work per item and is considerably less work overall, because it removes the combination problem entirely.
How many words do you actually need
Fewer than most people assume for conversation, and more than they hope for comfort.
A few thousand well-controlled words cover the large majority of everyday speech, because word frequency is steeply skewed — a small set does most of the work. What separates a learner who manages from one who struggles is usually not vocabulary size but whether those words are active or passive.
This is why chasing word counts is a poor target. Two thousand words you can deploy under time pressure will serve you better than eight thousand you can recognise. The CEFR descriptors reflect this: they describe levels by what you can do, not by how many items you have accumulated, and the Europass grid rates speaking separately from reading — which is where the active-passive gap becomes visible.
Related reading
- why you understand but cannot speak
- how to build spoken fluency with an AI tutor
- understanding CEFR levels and your speaking score
- the shadowing technique explained
- our 2026 ranking of the major AI language apps
- AI language learning vs Duolingo