Before anything else: we build one of these. OwnSlides is a language-learning platform, so you should read what follows with the same suspicion I am about to aim at everyone else’s research.
With that out of the way. The evidence here comes from two directions, both compromised. There are studies showing impressive results, frequently funded by the companies whose products are being tested, and flat declarations that apps are useless, resting on somebody’s abandoned 400-day streak. Neither is much use. Sources are linked inline so you can check every claim yourself.
What the company-funded studies show
Several app makers have published or commissioned research on their own products. The recurring headline is that X hours on the app matches a semester of university study, usually measured with standardised reading and listening tests.
Dismissing these purely on funding is lazy, since many use real instruments and report real gains. But the same limitations keep turning up:
- Self-selected, motivated participants. People who volunteer for a study about an app they already like are not the median downloader, who quits in week two.
- Outcomes favour what apps train. Multiple-choice reading and listening tests measure recognition, which is exactly what app drills build and exactly what overstates your conversational ability.
- Comparison to university courses is flattering. A beginner university course is a low bar on hours-to-outcome, since a lot of class time goes on administration and waiting your turn.
What independent evidence suggests
Away from company money the picture is more measured, though not negative. Two decades of research on computer- and mobile-assisted language learning broadly finds that well-designed tools produce real gains, particularly for vocabulary, which makes sense, since spaced self-testing is precisely what the memory literature endorses.
The consistent qualifier is that gains concentrate in some skills and not others, and that engagement collapses. Retention curves for language apps are brutal. An app can work perfectly well for the people who persist and still fail the overwhelming majority who download it.
The most important variable in whether an app works is not its teaching method. It is whether you are still opening it in month four.
What apps are actually good at
Specifics beat verdicts. Apps are strong at:
- Vocabulary through spaced retrieval. The best-evidenced thing they do. Software schedules reviews better than you ever will, and it doesn’t get bored.
- Building a daily habit. Streaks get mocked relentlessly and they work. They target the one variable that dominates outcomes, which is whether you came back.
- Low-stakes early practice. You can be wrong in private. That matters more than it sounds for people who would otherwise never start.
- Structured beginner progression. Sequencing material sensibly for a novice is real pedagogical work, and the good ones do it well.
What they systematically miss
And equally specifically, the things they are bad at, ours included:
- Spontaneous production. Tapping word tiles into the right order is not composing a sentence while a human being waits for you to finish. This is the most common complaint from long-term app users and it’s justified.
- Unscripted listening. App audio is clean, slow and studio-recorded. Real speech is fast and full of swallowed syllables. The transfer is much weaker than your app scores suggest.
- Recovering from breakdown. Real conversation is largely misunderstanding and recovery. Apps rarely simulate it, and it’s a skill of its own.
- The intermediate plateau. Most apps are strong to about A2 and thin out after B1. That upper-intermediate stretch, where progress slows and good material gets scarce, is where people quit.
The honest verdict
Apps work for what they are: an efficient way to build vocabulary, a daily habit and a beginner foundation. They don’t work as a complete route to conversational fluency, and the gap gets wider the further you go.
The people who get the most out of them treat the app as one component and go looking elsewhere for what it can’t give them: real listening material, actual conversation, and enough grammar to decode what they’re hearing.
Which suggests a sensible way to use one:
- Use it for vocabulary and daily consistency, which is where the evidence is strongest.
- Add unscripted listening early, even when it is uncomfortably hard.
- Start speaking well before you feel ready; this is the gap apps leave widest.
- Expect to outgrow the app around B1 and plan what replaces it.
We built our own platform around those specific gaps, with conversation practice and unscripted listening sitting alongside the vocabulary work rather than after it. Whether that closes the gap is for you to judge, and no app fully replaces talking to a person.
The question isn’t whether apps work. It’s which part of the job you’ve handed to one, and where you’re getting the rest.



