Skip to main content

Cartesia vs ElevenLabs for voice agents

Both are very good, and they are good at different things. Cartesia is our default text-to-speech; ElevenLabs voices ship in the voice library. This is what the choice actually looks like once the calls are real, written by people who run both rather than people who read both changelogs.

The trade

The trade, stated plainly

ElevenLabs built its reputation on voice realism and a very large multilingual voice library, and that reputation is deserved — for narration, content and anywhere a human listens to a long uninterrupted passage, it is exceptional.

Cartesia is built around streaming for real-time conversation, which is a different optimisation target. A phone call does not want the best possible rendering of a paragraph; it wants the first syllable to arrive before the silence gets awkward, and it wants synthesis to stop instantly when the caller interrupts.

That is why Cartesia is our default and why the answer is not universal. The right question is not which engine is better but which failure you would rather have: a slightly less remarkable voice, or a slightly longer pause.

Latency

What matters on a phone call

Three things, roughly in this order, and only one of them is what demos are judged on.

Time to first audio byte. Not total synthesis time. The caller hears the beginning of the sentence while the end is still being generated, so the number that matters is when audio starts. Ask for it under streaming conditions and on your own text.

Interruption behaviour. When someone talks over the agent, synthesis has to stop — properly, not after finishing the current chunk. An agent that keeps talking for another second and a half feels worse than one with a slightly duller voice. Barge-in is a requirement, not a feature.

How it handles your vocabulary. Every business has words that break TTS — surnames, drug names, SKUs, street names, your own product. Both engines can be corrected; test them on your actual terms rather than on a paragraph of generic prose, because the failure will be specific to you.

Voice quality

Where the phone call constrains both

Worth knowing before agonising over the choice: the telephone network resamples audio to 8 kHz. A great deal of what separates two premium engines in a side-by-side listen on good headphones is simply not present by the time it reaches a caller’s handset.

This is the most common way teams waste a fortnight. The differences are real in a studio and compressed on a phone line, while latency and interruption handling are not compressed at all — they are exactly as noticeable at 8 kHz. Spend the fortnight on the latency budget instead; the gains there are larger and they survive the codec.

Choosing

Common questions

Which one is faster?
We have not published a controlled benchmark between them and will not quote one we did not run. What we can say is architectural: time to first audio byte matters far more than total synthesis time on a phone call, because the caller hears the start of the sentence while the rest is still being generated. Ask any vendor for time-to-first-byte under streaming, not for characters per second, and test it on your own text — numbers on marketing pages are rarely measured the way a phone call works.
Can we use both?
Yes, and there is a sensible reason to. TTS choice is per-agent rather than per-account, so a high-volume outbound campaign and a flagship inbound line do not have to make the same trade. Cartesia is the default; ElevenLabs voices are available in the voice library.
Does the voice actually change conversion?
It changes how long people stay on the call, which is upstream of everything else. But the effect is smaller than teams expect and it saturates quickly — past a certain quality bar, callers respond to whether the agent understood them and did the thing, not to timbre. Latency and interruption handling are usually worth more attention than voice selection, and they are cheaper to fix.

No benchmark figures appear on this page. We have not run a controlled latency or quality comparison between these two engines, and quoting one we did not run would undermine the only thing that makes this page worth reading.

Hear both on your own script

Bring the words your business actually says — the surnames, the product names, the street. That is where the two diverge.