What the TTS layer does
TTS turns the language model's text output into audio the caller hears. It is one of three AI services on every Finn, alongside STT and the LLM.
You do not usually pick a TTS engine directly. You pick a voice, and the voice determines the engine. Voice selection is Voice persona, under Call flow → Identity in the editor (or on the Voice & call behavior tab when you build from a playbook). See Voice for the voice library itself.
The default
Every Finn is created with a default TTS service already set. It is tied to whichever voice you selected. The editor has no separate TTS field: picking a voice sets it. Only the playbook setup shows a TTS service field, in the Advanced mode of its Voice & call behavior tab.
That field is hidden by default for a reason: changing TTS independently of voice is not something most accounts need. The documented guidance is to change it only if you have a specific reason.
What actually changes the audio
Three fields, in order of how much they matter.
| Field | Where | Effect on audio |
|---|---|---|
| Voice | Call flow → Identity → Voice persona | Determines timbre, pacing, gender, accent, and which TTS engine runs |
| Language | Call flow → Identity → Default language | Filters the voice library to voices that support it, and constrains TTS |
| TTS service | Playbook setup only: Voice & call behavior → Advanced | The engine behind the voice. Tied to voice selection |
Changing Language filters the voice list. If your current voice does not support the new language, you have to reselect. Not every voice supports every language.
Latency
TTS is part of the response latency the caller experiences, and engines differ. Test a voice on a real call before committing to it on a latency-sensitive use case.
Above roughly 1.5s the conversation stops feeling natural. Callers start talking over the agent, which then feeds barge-in into STT and compounds the problem. If your Finn feels sluggish, the voice you picked is the first thing to check, before you start editing prompts.
When to change it
Change the voice, not the TTS service, in almost every case. Reasons to change voice:
- The tone does not match the use case. Calm and warm for healthcare, clear and direct for logistics, energetic for sales.
- You switched Language and the current voice does not cover it.
- Latency is too high and you are on a premium voice.
- Callers report the voice is hard to understand.
Reasons to touch the TTS service field itself are narrow, and if you cannot state yours in one sentence, leave it alone.
What TTS will not fix
TTS controls how words are spoken, not which words. If the agent says the wrong thing, that is the LLM and your prompt, not the voice. See Prompting.
Two common misattributions:
| Symptom | Actual layer |
|---|---|
| Agent misheard the caller | STT, see STT |
| Agent rambles or gives long answers | Response Guidelines in the prompt, see Prompting |
Multilingual behavior
A Finn has one default language, and its voice has to support that language. There is no auto-detect mode that switches language mid-call. If you change the language while building from a playbook, the personality step offers to translate the personality fields for you. Either way, you may need to reselect the voice. See Multilingual.
Dashboard, not API
Voice, Language, and TTS service are set in the Finn editor in the dashboard. Through the API you can set voice and language, but not the TTS service: the API rejects any provider field, and setting voice puts the Finn on that voice's provider. The Voices API lists every voice your organization can use. See api-finns.
Changes apply to new calls. Calls already in flight finish on the old configuration, so a voice change during a live campaign produces a mixed set of recordings. Test the new voice on yourself first with Test → Phone call, which routes a real call to your phone. See Testing.
Next
- Voice — the voice library and how to pick.
- STT — the input side of the audio path.
- Multilingual — language settings.