Every voice AI vendor is shipping fast, and every one of them wants you to know it. Open any competitor blog this quarter and you'll find a "2025 Product Recap," an "Everything We Shipped in August," or a changelog note celebrating that Spanish and German speech-to-text now format numbers correctly. Useful for the vendor's marketing calendar. Useless for a buyer trying to decide where to put a six-figure contract.
The recap format hides the one thing you actually need: which of these updates is now baseline, and which is a real differentiator you should pay for? When a feature shows up in four changelogs in one quarter, it's not an edge anymore — it's the floor. Confusing the two is how teams overpay for "Salesforce integration" that every vendor has, while under-scoping the latency work that actually decides whether the agent survives a live call.
This is the buyer's read. We'll walk the 2026 update cycle by capability, sort table-stakes from differentiator, give you the questions to ask any vendor, and show where Finn sits on each axis — without the self-recap tone.
Why competitor "product recap" posts don't help buyers
A product recap answers "what did we do." A buyer needs "what does everyone do, and where's the gap." Those are different documents.
Three problems with the recap genre:
- No baseline. A changelog that says "we added HubSpot sync" implies novelty. But if Retell, Vapi, Bland, and Finn all added CRM sync the same year, HubSpot sync is table-stakes — celebrating it is noise. You can't tell from a single vendor's post.
- Feature name ≠ feature depth. "Salesforce integration" can mean a one-way webhook that logs a call, or a bi-directional sync that reads open opportunities mid-call and writes disposition + next-step back to the record. Same two words, 10x difference in value.
- Shipping cadence gets sold as quality. "We shipped 47 features this quarter" tells you about the vendor's marketing, not whether barge-in works at 300ms. Cadence is a proxy, and a weak one.
The fix is to read changelogs horizontally — across vendors, by capability — instead of vertically down one company's timeline. That's what the rest of this piece does.
Integrations that became table-stakes in 2026 (Salesforce, Calendly, CRM)
By mid-2026, the following stopped being differentiators and became the price of entry. If a vendor treats any of these as a headline feature, they're behind:
- CRM write-back — Salesforce, HubSpot, GoHighLevel. Not "we can POST to a webhook," but native objects: log the call, update the contact, set disposition, create a follow-up task.
- Calendar booking — Calendly, Google Calendar, Cal.com. The agent checks real availability and books inside the call, no callback.
- Telephony bring-your-own — Twilio, Vonage, SIP trunk. You keep your numbers and carrier relationship.
- Post-call export — transcript, recording, structured summary pushed to your warehouse or ticketing tool.
Here's the buyer test that separates real integration from a checkbox: is it bi-directional and mid-call, or one-way and post-call?
A one-way "Salesforce integration" logs the call after it ends. A real one lets the agent read the CRM during the call — "I see your last order shipped Tuesday, is that what you're calling about?" — and write structured fields back the instant the call closes. The first is a log. The second changes the conversation. Most changelogs won't tell you which one they built. Ask.
The same test applies to salesforce calendly integration bundles vendors love to advertise: booking that reads live availability and writes the event is table-stakes; booking that emails a link is not integration, it's a fallback.
Omnichannel: SMS + web widget + voice as one context (not bolted on)
The 2026 buzzword is omnichannel customer solutions, and it's where the "we shipped it" gap is widest. Nearly every platform added sms and web widget support this year. Almost none unified the context behind those channels.
The distinction that matters:
- Bolted-on omnichannel: SMS, web chat, and voice are three separate agents with three separate memories. Customer texts, then calls — the voice agent knows nothing about the text.
- Unified-context omnichannel: one conversation state across channels. Customer starts on the web widget, escalates to a call, and the voice agent picks up mid-thread with full history.
The first is three products in a trench coat. The second is an enterprise voice platform. The tell in a changelog: does the SMS release share a session/context store with voice, or is it a standalone module with its own dashboard? If the vendor ships them as separate "products" with separate pricing, they're bolted on.
Buyer question: "If a customer texts us, abandons, and calls back an hour later, does the voice agent see the SMS thread?" If the answer needs a Zapier diagram, it's not unified.
Multilingual & universal speech-to-text — Spanish, German, beyond
This is the category most abused by recap posts. "We improved spanish speech to text." "German number formatting fixed." Real work — but framed as a feature when it's really the platform catching up to universal speech to text as the new default.
In 2026, multilingual transcription is table-stakes for anyone selling outside a single English-speaking market. What still varies — and what you should actually probe:
- Code-switching — real callers mix languages mid-sentence ("necesito un refund for order 1-2-3"). Does the STT handle it, or does it reset to one language and drop the rest?
- Accent and dialect coverage within a language —
german speech to textthat nails Hochdeutsch and falls apart on Austrian or Swiss German isn't "German support." - Domain vocabulary — drug names, SKUs, policy numbers. Universal STT still fumbles proper nouns without a custom vocabulary hook.
- Latency parity — some engines add 200–400ms for non-English. If your Spanish calls lag your English calls, that's a differentiator hiding in a "we support Spanish" line.
Baseline: "supports Spanish and German." Differentiator: code-switching, custom vocabulary, and equal latency across languages. If you build in the STT layer directly, our Google Speech-to-Text builder's guide breaks down where universal models still need help.
What's still a real differentiator (latency, barge-in, grounding)
Strip away the integrations everyone shipped and three things still separate a demo from a production agent — and none of them photograph well in a changelog:
- Turn latency. End-to-end response under ~800ms feels like a conversation; 1.5s feels like a hold. This is the hardest thing to fake and the first thing that breaks under load. Ask for p95 latency under concurrent calls, not a demo number.
- Barge-in. Can the caller interrupt the agent mid-sentence and be understood immediately? Real humans interrupt. An agent that talks over you or ignores the interrupt loses trust in one turn.
- Grounding. Does the agent answer from your knowledge base and live systems, or does it improvise? A grounded agent says "I don't have that, let me transfer you." An ungrounded one hallucinates a refund policy. This is where retrieval quality and tool-calling reliability matter more than model choice.
These are the axes that don't show up as one-line changelog wins because they're systems work, not features. That's exactly why they still differentiate. For the deeper build view, see what separates production voice agents from no-code prototypes.
The questions to ask any voice AI vendor before you sign
Print this. Walk it through every demo:
- Integrations: "Is your Salesforce/HubSpot integration bi-directional and mid-call, or post-call logging only?"
- Omnichannel: "Do SMS, web, and voice share one conversation context, or separate sessions?"
- STT: "How do you handle code-switching and custom vocabulary? What's latency for Spanish/German vs English?"
- Latency: "What's your p95 turn latency at 100 concurrent calls, not in a solo demo?"
- Barge-in: "Show me an interruption mid-sentence, live."
- Grounding: "What does the agent do when it doesn't know the answer? Show me a wrong-question."
- Telephony: "Can I bring my own Twilio/SIP and keep my numbers?"
- Data: "Where do transcripts and recordings land, and can I export raw?"
Any vendor who answers all eight crisply has built a real platform. Anyone who redirects to their changelog hasn't.
Where Finn stands on each capability
No recap tone — just the map:
- CRM + calendar: Bi-directional, mid-call. Finn reads live CRM and availability during the call and writes disposition + booking back on hang-up. Table-stakes, done at the depth that counts.
- Omnichannel: SMS, web widget, and voice share one conversation context. A caller who started on chat is picked up mid-thread on the phone.
- Multilingual STT: Universal speech-to-text with code-switching and custom vocabulary; latency parity across supported languages including Spanish and German.
- Differentiators: Sub-800ms turn latency under concurrent load, native barge-in, and grounded responses tied to your knowledge base with explicit "I don't know → transfer" behavior.
We'd rather you run the eight questions against us than take a recap's word for it — including our competitors'. Compare on the axes that decide live calls: latency, grounding, and whether the integration actually talks back.
FAQ
What are the must-have voice agent integrations in 2026? Bi-directional CRM write-back (Salesforce, HubSpot), live calendar booking (Calendly, Google Calendar), bring-your-own telephony (Twilio, SIP), and post-call export. These are table-stakes — if a vendor headlines them as new, they're behind.
Is multilingual speech-to-text still a differentiator? Basic Spanish/German support is now baseline. The real differentiators are code-switching (mixed languages mid-sentence), custom domain vocabulary, and equal latency across languages.
What separates a production voice agent from a demo? Turn latency under load (~800ms p95 at concurrency), barge-in that handles interruptions, and grounding that keeps the agent from hallucinating. None photograph well in a changelog, which is why they still differentiate.
How do I evaluate omnichannel claims? Ask whether SMS, web, and voice share one conversation context or run as separate sessions. If a customer texts then calls, the voice agent should see the text thread without a Zapier workaround.
(Emit FAQ JSON-LD from the four Q&As above.)
Related reading
- AI Voice Agent Pricing Comparison 2026
- AI Voice Agent vs IVR — Enterprise Guide
- Google Speech-to-Text API: The 2026 Builder's Guide
- Voiceflow Alternatives for Production Voice Agents
CTA
Run the eight questions against Finn. Book a live demo and interrupt the agent, ask it something it shouldn't know, and watch it read and write your CRM mid-call — no recap required. See Finn on a real call →



