Speech technology
Text-to-speechTTS
Also called: speech synthesis · voice synthesis
Text-to-speech (TTS) converts written text into spoken audio. In a voice agent, TTS is the final step — it speaks the agent's response aloud in a natural voice, ideally in the caller's language and with low latency.
Modern neural TTS sounds close to human, with natural intonation and the ability to clone or customize voices. Quality and speed matter: robotic or laggy TTS breaks the illusion of a real conversation.
See it in the product
Related terms
Speech-to-textTechnology that transcribes spoken audio into text in real time. STT is how a voice agent understands what the caller is saying.Voice cloningCreating a synthetic text-to-speech voice that mimics a specific person's voice from a short sample. Used to give agents a consistent brand voice.Voice latencyThe delay between a caller finishing speaking and the agent responding. Low latency (sub-second) is what makes an AI voice conversation feel natural.Voice AIArtificial intelligence applied to spoken language — recognizing speech, understanding intent, and responding with a synthetic voice in real time.
See Text-to-speech in a real call.
Book a 30-minute demo and watch Finn handle inbound and outbound calls end to end — no stack to assemble.
Try Finn for free