Speech technology
Automatic speech recognitionASR
Also called: speech recognition
Automatic speech recognition (ASR) is the technology that converts spoken language into text. ASR and speech-to-text (STT) are the same capability — ASR is the academic/engineering term, STT the product term.
ASR quality is measured by word error rate (WER). For voice agents, streaming ASR (transcribing incrementally as the person speaks) matters more than batch accuracy, because it enables fast turn-taking.
Related terms
Speech-to-textTechnology that transcribes spoken audio into text in real time. STT is how a voice agent understands what the caller is saying.Natural language understandingThe part of AI that works out what a person means — their intent and the key details — from natural language, not just the literal words.EndpointingDetecting when the caller has finished speaking so the agent knows it's its turn to respond. Bad endpointing causes awkward pauses or interruptions.
See Automatic speech recognition in a real call.
Book a 30-minute demo and watch Finn handle inbound and outbound calls end to end — no stack to assemble.
Try Finn for free