What the language model does
The language model is the reasoning layer of a Finn. It reads the conversation so far — your identity text, style and guardrails, response guidelines, retrieved knowledge base snippets, and every turn of the call — and decides what the agent says next and when to call a tool.
It is one of three AI services behind a call. The other two are text-to-speech, which turns the model's output into audio, and speech-to-text, which turns the caller's audio into text. See text to speech and speech to text.
Where you change it
In the dashboard. In the editor, it is Call setup → Speech settings → Language model. When you build from a playbook, switch the setup's Voice & call behavior tab to Advanced; the field is labelled LLM service.
Model choice is not available through the API. The finn API does not accept a model or provider field: providers are set per organization.
Every Finn ships with a default model that works for the large majority of use cases. Change it only when you have a specific reason and a way to measure whether the change helped.
What a model change actually affects
| Area | Effect of switching models |
|---|---|
| Instruction adherence | How reliably the agent honours "never" rules in guardrails |
| Latency | Time between the caller finishing a sentence and audio starting |
| Tool calling | Whether tools fire at the right moment with the right arguments |
| Knowledge base use | How well retrieved snippets are woven into an answer rather than read aloud verbatim |
| Cost per call | Model tier is a per-minute cost input |
The model does not affect how the agent sounds. Accent, pace, and warmth come from the voice, not the model. If callers say the agent sounds robotic, the fix is in voice or text to speech, not here.
Latency is the real trade-off
Voice is unforgiving about delay in a way that chat is not. A caller hears a pause and starts talking again, which produces an interruption the agent then has to recover from. A model that produces a better answer 400 ms later can measurably degrade a call.
Prompt length compounds this. See the prompting guide.
If your agent is slow, shorten the prompt before you change the model. A 2000-word prompt on a fast model is usually slower than a 400-word prompt on a slower one, and the shorter prompt is also more reliable.
Diagnose before you switch
Most behaviour problems blamed on the model are missing-rule bugs. Before changing the LLM service, work through this:
- Listen to the recording in call logs. Do not guess from the transcript summary.
- Find the first turn where the conversation went wrong. Every later problem usually descends from it.
- Check whether a rule covering that exact situation exists in your guardrails or response guidelines. Usually it does not.
- Add the rule and one example, then re-test.
If the same failure survives a specific, explicit rule plus an example, that is a genuine model-capability signal and worth acting on. If you have not written the rule yet, a model change is guesswork.
Testing a change
A model change is a configuration change like any other, so the safety rules apply. It affects new calls only. In-flight calls finish under the old configuration.
Test with Simulation chat first, in the browser, to check reasoning and tool calls without placing a call. Then test with a Phone call, because latency and interruption behaviour only show up in real audio. See testing.
Run the same set of scenarios you ran on the old model, including the hostile caller, the confused caller, and the caller who goes off-topic. A model that is better on the happy path can be worse at recovery, and recovery is where calls are lost.
There is no configuration history to roll back to. Note the model you are switching from before you change it, so you can switch back if the change degrades performance.
Measuring the result
Do not judge a model change on a handful of test calls. Deploy it, let real volume accumulate, and compare structured outputs before and after. Post-call analysis fields such as appointment_booked or sentiment give you a comparable number across two configurations. Evaluations and call outcomes cover this in detail.
Which specific models are offered is shown in the Language model dropdown in the dashboard. That list changes, so read it there rather than assuming a model you used previously is still the one selected.