Where testing lives
Open a Finn from Voice AI Agents in the sidebar. The Test menu in the top right has three options:
| Option | What happens | Use it for |
|---|---|---|
| Phone call | The Finn calls your phone | Voice, pacing, interruptions, how the welcome message actually lands |
| Simulation chat | Text simulation in the browser | Prompt logic, knowledge base retrieval, fast iteration |
| Web call | A voice call in the browser | Hearing the voice without using a phone |
The builder's Simulation test tab goes further. Besides manual chat, it has a Test suite mode, where you save conversations with assertions and run them all at once, and an Input noise test mode that generates variants such as STT noise, paraphrases, and persona shifts.
All of these are dashboard actions. There is no API for the Test menu itself, but POST /calls places a real call with any Finn, draft or live, so you can script a call to your own number. See api-calls.
Test phone calls are not free. They need credits in your wallet, and a call is refused with "Insufficient credits" if the balance is empty. A test call is also refused while that Finn has a scheduled deployment that has not started yet. Nothing you do here is live to the outside world. A Finn only reaches real callers once you create a deployment, covered in outbound campaigns and inbound routing.
Chat first, call second
Chat is the faster loop. It exercises the same identity text, guardrails, response guidelines, and knowledge base, without the latency of a phone call. Use it to settle what the Finn says.
Chat does not exercise the voice, TTS, or STT stack. It will not surface a mispronounced brand name, a voice that sounds wrong for the use case, an idle timeout that fires too early, or the Finn talking over the caller. Those only show up on a real call. See voice, tts, and stt.
Move to a phone call once the script is roughly right, then keep alternating.
Phone number format
Enter the number with a country code. US: +15551234567. India: +919876543210. The dashboard shows a preview of the format it parsed before it dials. Read that preview. A number without a country code is the most common reason a test call never arrives.
What to test, not just that it works
A test that only covers a cooperative caller answering in order tells you little. Cover:
- Callers who go off-topic or ask something outside the knowledge base
- Callers who are confused and ask the same thing twice
- Hostile or impatient callers, to check the guardrails hold
- Accents your speech recognition may not handle, and brand names or SKUs it may mishear
- Silence, to see when the idle reminder fires and what it says
- Answers in an unexpected format, for example a date with no time
Reading the result
Every test call lands in Analytics alongside real calls. Open the call to get the transcript, the audio, and the side panel. Clicking a line in the transcript jumps the audio to that moment.
When the Finn answers a question wrongly and you believe the answer is in your documents, check the knowledge base content before editing the prompt. See knowledge base and call logs.
Things that will go wrong
| Symptom | Cause | Fix |
|---|---|---|
| Call never arrives | Missing country code | Re-enter with + and country code |
| Call refused before dialing | Wallet empty, or a scheduled deployment is pending for this Finn | Add credits, or wait until the scheduled deployment has run |
| Call never arrives | Carrier or handset spam filter dropped it | Try a different from-number |
| Finn sounds wrong or fails to speak | Voice not assigned, or voice needs a region match | Re-select the voice |
| Post-call analysis fields blank | The field was added after this call ran | Add the field, then make a new test call |
| Post-call analysis fields blank | Call was too short to extract from | Run a longer test |
| Finn ignores your instructions | Contradictions across Identity, Style, Welcome, and Workflow | Audit all four for conflicts. See prompting and guardrails |
Post-call analysis is the trap most people hit. It runs on calls that complete after the field exists, never retroactively. Define your fields before a long test session, not after.
Edits during testing
Saving a Finn does not interrupt anything. Changes apply to new calls. A call already in flight finishes under the configuration it started with, so a config edit mid-call will not show up until you dial again.
There is no configuration history or rollback. Before a significant edit, copy the current prompt or duplicate the Finn, so you are not reconstructing the earlier version from memory.
When testing stops being enough
Manual testing catches obvious breakage. It does not tell you whether a change improved behavior across many calls, because you cannot hold twenty transcripts in your head. Once the Finn is roughly right, save your key conversations as a test suite, and move to evaluations for repeatable scoring and post-call analysis for measurement across a real call volume.