What speech to text does
Speech to text (STT) turns the caller's audio into text the LLM can act on. Every turn of a call runs through it. If STT hears the wrong thing, the LLM answers the wrong question, the workflow branches wrong, and post-call analysis extracts wrong values. Most "the Finn didn't follow my script" reports trace back to a transcription miss, not a prompting problem.
STT is configured on the Finn, in the dashboard. The API does not accept an STT service: providers are set per organization. See api-finns, and agents-identity for where language is set.
Fields that control STT
| Field | Where | What it does |
|---|---|---|
| Language | Call flow → Identity → Default language | Primary recognition language. Also drives transcript output and filters the voice library. |
| STT service | Call setup → Speech settings → Speech to text (Voice & call behavior → Advanced in the playbook setup) | Which speech-to-text provider runs. Default works for most languages. |
| Speech model | Realtime transcription, in the Edit modal | The Deepgram model used for transcription. |
| Endpointing (ms) | Realtime transcription, in the Edit modal | How long a pause counts as the end of the caller's turn. |
Language is the field that matters most. Pick what your callers actually speak, not what your website is in. Setting a Finn to English and pointing it at Tamil-speaking callers does not degrade gracefully — it produces confident nonsense transcripts rather than empty ones.
Accuracy by language
The language picker offers 47 languages, and STT quality is not uniform across them. For high-stakes use cases (sales, healthcare, finance), test with real callers in your target language before launch.
Accents
English recognition is robust to most accents, but not immune. Heavy regional accents produce higher error rates, and there is no dashboard readout that tells you your error rate for a given accent group.
The only way to know is to test. Run test calls with speakers who match your real caller base, then read the transcripts in call-logs and compare against what was actually said. Do this before launching, not after.
Things that make accent handling worse, in rough order of impact:
- Poor audio (speakerphone, hands-free car kits, noisy call centers)
- Short single-word answers with no surrounding context to disambiguate
Code-switching
Callers in many regions mix languages inside one sentence — Hindi-English, Spanish-English, French-Arabic. To stop the Finn from responding in a single fixed language, add to the system prompt:
Match the caller's speaking style — if they mix English and Hindi, you can mix too. Don't correct or comment on their language choice.
Code-switching is not reliable for every language pair. If your callers code-switch heavily, test with real mixed-language speech and expect some transcript errors in the mixed spans.
What goes wrong
| Symptom | Likely cause | Fix |
|---|---|---|
| Finn answers a question the caller did not ask | Transcription miss upstream | Read the transcript in call logs to find where the caller was misheard |
| Finn asks the same question twice | Caller's answer was misheard or cut off | Read the transcript, then rephrase the question or raise Endpointing. See agents-conversation |
| Accuracy dropped after a config change | STT service or speech model changed | Revert the change and compare transcripts |
| Post-call analysis fields blank or wrong | Transcript too short or too corrupted to extract from | Check transcript length first. See post-call-analysis |
| Long silences before the Finn responds | Endpointing set too high, or idle timing | Lower Endpointing, then check Reminder message frequency. See agents-conversation |
Related
- agents-multilingual — language settings
- agents-tts — the output half of the audio path
- agents-testing — test calls before you launch
- troubleshooting — broader symptom index