Skip to main content

Agents

Speech to text

Transcription settings, languages and accent handling.

What speech to text does

Speech to text (STT) turns the caller's audio into text the LLM can act on. Every turn of a call runs through it. If STT hears the wrong thing, the LLM answers the wrong question, the workflow branches wrong, and post-call analysis extracts wrong values. Most "the Finn didn't follow my script" reports trace back to a transcription miss, not a prompting problem.

STT is configured on the Finn, in the dashboard. The API does not accept an STT service: providers are set per organization. See api-finns, and agents-identity for where language is set.

Fields that control STT

FieldWhereWhat it does
LanguageCall flow → Identity → Default languagePrimary recognition language. Also drives transcript output and filters the voice library.
STT serviceCall setup → Speech settings → Speech to text (Voice & call behavior → Advanced in the playbook setup)Which speech-to-text provider runs. Default works for most languages.
Speech modelRealtime transcription, in the Edit modalThe Deepgram model used for transcription.
Endpointing (ms)Realtime transcription, in the Edit modalHow long a pause counts as the end of the caller's turn.

Language is the field that matters most. Pick what your callers actually speak, not what your website is in. Setting a Finn to English and pointing it at Tamil-speaking callers does not degrade gracefully — it produces confident nonsense transcripts rather than empty ones.

Accuracy by language

The language picker offers 47 languages, and STT quality is not uniform across them. For high-stakes use cases (sales, healthcare, finance), test with real callers in your target language before launch.

Accents

English recognition is robust to most accents, but not immune. Heavy regional accents produce higher error rates, and there is no dashboard readout that tells you your error rate for a given accent group.

The only way to know is to test. Run test calls with speakers who match your real caller base, then read the transcripts in call-logs and compare against what was actually said. Do this before launching, not after.

Things that make accent handling worse, in rough order of impact:

  • Poor audio (speakerphone, hands-free car kits, noisy call centers)
  • Short single-word answers with no surrounding context to disambiguate

Code-switching

Callers in many regions mix languages inside one sentence — Hindi-English, Spanish-English, French-Arabic. To stop the Finn from responding in a single fixed language, add to the system prompt:

Match the caller's speaking style — if they mix English and Hindi, you can mix too. Don't correct or comment on their language choice.

Code-switching is not reliable for every language pair. If your callers code-switch heavily, test with real mixed-language speech and expect some transcript errors in the mixed spans.

What goes wrong

SymptomLikely causeFix
Finn answers a question the caller did not askTranscription miss upstreamRead the transcript in call logs to find where the caller was misheard
Finn asks the same question twiceCaller's answer was misheard or cut offRead the transcript, then rephrase the question or raise Endpointing. See agents-conversation
Accuracy dropped after a config changeSTT service or speech model changedRevert the change and compare transcripts
Post-call analysis fields blank or wrongTranscript too short or too corrupted to extract fromCheck transcript length first. See post-call-analysis
Long silences before the Finn respondsEndpointing set too high, or idle timingLower Endpointing, then check Reminder message frequency. See agents-conversation