Why we made Litani voice-first (and why tapping the right answer isn't speaking)
Recognition feels like learning. Production is the thing you actually need.
Most language apps ask you to point at the right answer. Four tiles appear, you tap the one that matches, a chime plays, and a little number goes up. It feels like learning. You are getting things right.
Then you sit down across from a real person, they ask you something ordinary, and nothing comes out.
That gap is the whole reason Litani is voice-first. Tapping the right answer and saying it are not the same skill, and only one of them is the one you actually want.
What is the difference between recognition and production?
Recognition is picking the correct option out of a set. You see the word, you match it, you move on. The answer was sitting right in front of you the whole time.
Production is making the words yourself, from nothing, with no menu to choose from. Someone asks where the station is and you have to assemble the sentence, find the verb, and get the sounds out while they wait.
Recognition is much easier, which is exactly why it feels good. You almost always get it right, so you feel like you are making progress. But speaking in real life is pure production. There is no menu. The two skills overlap less than you would hope, and practising the easy one does not quietly train the hard one.
The short version
If your practice never asks you to speak, you are getting very good at the one thing you will never have to do in a conversation: choosing from four options.
Why does saying it out loud work better?
There is a well-studied memory finding called the production effect. Words you say aloud while studying are remembered better than words you only read silently. Saying something gives it a more distinctive trace in memory, so it is easier to find later.
The original work by Colin MacLeod and colleagues laid this out across a series of experiments. You can read their paper on the production effect if you want the detail.
For a language learner the implication is plain. Reading a translation is input. Saying it is output. Output is harder, and the difficulty is the point. The effort of producing the words is what makes them stick and what makes them available when you need to speak.
What does voice-first actually mean in Litani?
It means your input is your voice, not your finger. Here is the loop.
You start by capturing a real moment from your own life, about sixty seconds of it, spoken in plain English. Litani turns that into your practice material. Then, when you practise, you speak. Whisper transcribes what you said, and you get feedback on each word: Spot on, Almost, or Try again. The word that tripped you is shown, so you are not left guessing what went wrong. You can retry just that word and have another go.
That last part matters. Most speaking practice gives you a single pass and a vague verdict. Being able to see the exact word that slipped, and say it again right then, is how you actually fix it.
How Recall and Speak use your voice
Recall shows you an English prompt and asks you to say the phrase out loud in your target language, from memory. It checks what you said as you go. No tiles, no multiple choice. You produce the sentence or you don't, and the feedback tells you where you landed.
Speak goes further. It is a live, two-way conversation with an AI partner, the closest thing to rehearsing the real situation before you are in it. You talk, it responds, you keep the thread going. This is production under mild pressure, which is the rehearsal that recognition practice can never give you. If freezing mid-sentence is the thing you dread, this is the practice built for it. We wrote more about that in how to stop freezing mid-conversation.
What can the feedback actually tell you?
Here is the honest part. Litani listens to what you say and checks it word by word. It is good at that, and it is not a human examiner.
It can tell you reliably when a word landed, when it clearly slipped, and when you were somewhere in between. It is excellent for catching the words you keep dropping and for giving you a fast, judgement-free nudge to try again. What it cannot do is grade the fine shades of your accent or rule on a borderline case the way a patient teacher would.
So treat the verdict as guidance, not a sentence. Almost does not mean you have failed. It means say it once more. The goal is not a perfect run. The goal is that the words come out when a real person is standing in front of you.
How to use it
Chase the retry, not the score. If a word keeps landing on Almost, say it three more times and move on. The repetition is doing the work, not the grade.
The wager behind all of this
We made a bet when we built Litani. The bet is that the reason intermediate learners freeze is not that they lack vocabulary. It is that they have spent years recognising the language and almost no time producing it. So we put your voice at the centre and built every mode around saying things, not selecting them.
It is the harder path. Speaking is more effort than tapping, and the early sessions can feel exposing. But it is the only path that practises the actual skill. If you want the bigger picture on how this fits together, see what makes Litani different, and if you would rather absorb your phrases while your hands are busy, Podcast mode turns them into something you can listen to.
You can start with your own voice right now. Speak one real moment from your week and practise it back, out loud, today.
Frequently asked questions
What does voice-first mean in a language app?
It means your main input is your voice, not tapping or matching tiles. In Litani you speak, it transcribes what you said, and you get per-word feedback you can retry.
Why isn't tapping the right answer enough to learn to speak?
Tapping is recognition, where the answer is already in front of you. Speaking is production, where you build the sentence yourself. They are different skills, and only production prepares you for real conversation.
What is the production effect?
It is a memory finding that words you say aloud are remembered better than words you only read silently, because saying them leaves a more distinctive memory trace.
How accurate is Litani's speaking feedback?
It reliably flags whether a word landed, slipped, or came close, and shows you the word that tripped you so you can retry it. It checks the words you produced, not your accent, so treat it as a nudge rather than a final verdict.
Which Litani modes use my voice?
Recall asks you to say a phrase from an English prompt and checks it word by word. Speak is a live two-way conversation with an AI partner. Both are built around producing the language out loud.
Turn your life into practice
Capture a real moment, say it out loud, and start speaking with confidence.
Try Litani free