Skip to main content

Observability

Evaluations

Testing an agent before it answers a real caller.

What evaluation means on Finn

Finn has no automated evaluation of phone calls and no API endpoint that grades an agent. The builder's Simulation test tab can save chat conversations with assertions as a Test suite and run them together, but that covers text only, not voice or tools (see agents-testing). Evaluating real calls means placing test calls yourself, reading the transcript and the extracted fields, and fixing what you find before the agent takes a real caller.

That makes the loop manual, so the discipline matters more. The sections below are the order to work in.

Before the first test call

Check these in the dashboard. Each one fails in a way that looks like a model problem but is not.

CheckWhereFailure if skipped
Welcome message says who is calling and whyEdit → Call flow → Welcome messageCallers hang up on the opening
Voice assigned, region matchesEdit → Call flow → Identity → Voice personaSilence or wrong-sounding output
Knowledge base documents uploaded without a Failed statusEdit → Agent handbook → Knowledge baseAgent answers "I don't know" to documented facts
Post-call analysis questions definedEdit → Post-call analysisFields stay blank for every call in this test round
Trial credit remainingSettings → Plan and billing → UsageTest call never dials

Post-call analysis is the one that bites hardest. Fields only populate on calls that complete after the field was added. If you add a field mid-test, every call you already placed stays blank forever, and no re-run backfills them. Define your fields first, then test. See post-call-analysis.

Placing the test call

Enter the number with a country code. US is +15551234567, India is +919876543210. The platform shows a "tested format: …" preview before it triggers, so read that preview rather than trusting the number you typed.

A test call that never arrives is usually not the agent:

CauseCheck
Carrier rejectionSome international carriers block new numbers. Try a different from-number.
Spam filter on your handsetApple and Google filters drop calls silently. Try another from-number.

Full symptom list in troubleshooting.

Reading the result

Listen to the recording if the call has one (recording is an add-on set per deployment). Otherwise read the full transcript, not the summary. Find the first turn that went wrong, not the worst one. Everything after the first bad turn is downstream noise.

Then open the Data Extractor Sidebar for that call:

TabWhat it tells you
OverviewThe recording (if on), the call summary and the full transcript.
AI Copilot AnalysisThe AI's read of the lead from this call.
Post-Call AnalysisThe value each of your post-call fields got, and a "Needs Review" flag on answers the analysis marked for review.

If the agent stated something wrong that your knowledge base covers correctly, check that the document uploaded without errors and does not contradict itself before you touch the prompt. If the knowledge base is right and clean, the bug is in the prompt. See agents-knowledge-base.

Fixing and re-testing

The loop from the prompting guide applies directly:

  1. Find the first weird turn.
  2. Add a rule covering that exact scenario, plus one example turn.
  3. Re-test on the same number.
  4. If it is still wrong, the cause is upstream in the workflow, the audience, or the knowledge base.

Most "the model is dumb" reports are missing-rule bugs. Details in agents-prompting.

What a test call cannot tell you

Be honest about the limits, because a clean test call is weak evidence.

SignalWhy it misleads
SentimentAn indicator, not a verdict. It misclassifies. If it is consistently wrong for your use case, add a custom post-call field with a specific question and use that instead.
Very short callsCalls under five seconds are treated as no-answer by some metrics and will not appear where you expect. Check the full call log.
One clean passYou tested the happy path. Test the angry caller, the ambiguous request, and the caller who says "remove me" and expects the call to end.
Your own reactionsYou know the script. A real caller does not, and will interrupt, mumble, and change topic.

Test the negative rules explicitly. If the prompt says never quote a price, ask for a price. If it says never book a Sunday, ask for Sunday. A rule that has never been provoked has never been verified.

After you go live

Testing does not stop at launch. Watch the first real calls in call-logs and on the Live Deployments page (see monitoring). The failure modes that only appear at volume, such as hangups from a from-number carriers label as spam (see spam-labeling) or an exhausted audience, will not show up in a single test call.