Speech technology
Speech-to-textSTT
Also called: automatic speech recognition · ASR · voice recognition
Speech-to-text (STT), also called automatic speech recognition (ASR), transcribes spoken audio into text. In a voice agent, STT is the first step — it turns the caller's speech into text the language model can act on, as they talk.
Real-time STT has to handle accents, background noise, crosstalk and code-switching between languages. Its accuracy and speed set the ceiling for the whole conversation: if the agent mishears, everything downstream is wrong.
See it in the product
Related terms
Automatic speech recognitionThe formal name for the technology that converts speech into text. ASR and speech-to-text (STT) refer to the same thing.Text-to-speechTechnology that converts written text into natural spoken audio. TTS is what gives an AI voice agent its voice.Natural language understandingThe part of AI that works out what a person means — their intent and the key details — from natural language, not just the literal words.EndpointingDetecting when the caller has finished speaking so the agent knows it's its turn to respond. Bad endpointing causes awkward pauses or interruptions.
See Speech-to-text in a real call.
Book a 30-minute demo and watch Finn handle inbound and outbound calls end to end — no stack to assemble.
Try Finn for free