Skip to main content

Speech technology

Automatic speech recognitionASR

Also called: speech recognition

Automatic speech recognition (ASR) is the technology that converts spoken language into text. ASR and speech-to-text (STT) are the same capability — ASR is the academic/engineering term, STT the product term.

ASR quality is measured by word error rate (WER). For voice agents, streaming ASR (transcribing incrementally as the person speaks) matters more than batch accuracy, because it enables fast turn-taking.

See Automatic speech recognition in a real call.

Book a 30-minute demo and watch Finn handle inbound and outbound calls end to end — no stack to assemble.