Speech technology
Endpointing
Also called: turn-taking · end-of-speech detection
Endpointing is detecting the moment a speaker has finished their turn, so the agent knows when to start responding. It governs turn-taking — the rhythm of a natural conversation.
Endpoint too early and the agent cuts the caller off; too late and it feels sluggish. Good endpointing balances speed against not interrupting someone who's just pausing mid-thought.
Related terms
Barge-inThe ability for a caller to interrupt the agent mid-sentence and have it stop and listen — essential for natural, non-frustrating conversation.Voice latencyThe delay between a caller finishing speaking and the agent responding. Low latency (sub-second) is what makes an AI voice conversation feel natural.Speech-to-textTechnology that transcribes spoken audio into text in real time. STT is how a voice agent understands what the caller is saying.
See Endpointing in a real call.
Book a 30-minute demo and watch Finn handle inbound and outbound calls end to end — no stack to assemble.
Try Finn for free