Skip to main content

Speech technology

Speech-to-textSTT

Also called: automatic speech recognition · ASR · voice recognition

Speech-to-text (STT), also called automatic speech recognition (ASR), transcribes spoken audio into text. In a voice agent, STT is the first step — it turns the caller's speech into text the language model can act on, as they talk.

Real-time STT has to handle accents, background noise, crosstalk and code-switching between languages. Its accuracy and speed set the ceiling for the whole conversation: if the agent mishears, everything downstream is wrong.

See it in the product

See Speech-to-text in a real call.

Book a 30-minute demo and watch Finn handle inbound and outbound calls end to end — no stack to assemble.