On-prem vs hosted vs cloud
Three arrangements, and the middle one causes most of the confusion. On-prem is your software on your hardware in your building: complete control, and every upgrade, failure and capacity decision is yours. Hosted is typically that same software on somebody else’s hardware — relocated, not re-architected, still versioned and still capped per instance. Cloud is a multi-tenant service where capacity is elastic and upgrades arrive without a maintenance window.
Vendors use “hosted” and “cloud” interchangeably and they are not the same purchase. The question that separates them: if call volume triples on Monday, does capacity follow automatically, or does someone provision it?
What you stop maintaining
This is the honest benefit, and it is an operational one rather than a caller-facing one. Telephony cards and media servers, the patching cycle, capacity planning for seasonal peaks, and the disaster-recovery site that exists to be tested and never used. All of that stops being yours.
What does not go away is the call flow. The tree still has to be designed, and a badly designed menu is exactly as frustrating from the cloud as it was from the basement — the deployment model has no opinion about whether option four makes sense.
Failover and uptime
Uptime figures are quoted for the platform, and the platform is not the whole path. A call traverses your carrier, the SIP trunk, the network between them and then the service. A number quoted against the last hop tells you about the last hop.
Ask what happens when it fails rather than how often. Where do calls go — a fallback number, a recorded message, a busy tone — and does that failover trigger automatically or does somebody have to notice first? The second is common and rarely disclosed.
Cost model
The shift is capital to operating expense, and it changes who feels the cost. On-prem is a large purchase then years of amortisation, so growth is nearly free until it is suddenly not, at the moment you exceed capacity. Cloud is per-channel or per-minute: growth costs proportionally and nothing is stranded.
For steady, predictable volume on-prem can still be cheaper on a spreadsheet — that is a real answer, not a concession. For volume that spikes, elasticity is worth more than the unit rate, because the alternative is buying for the peak and idling through the rest of the year.
Migration
Porting the numbers is the long pole, and it has a fixed cutover — plan the go-live date around it rather than around the build. Rebuild the tree rather than transcribing it: an IVR that has accreted branches for a decade contains options nobody has chosen in years, and a migration is the cheapest opportunity you will get to delete them.
If you are moving anyway, it is also the moment to ask whether the menu is still the right shape — see IVR versus an AI voice agent. Not because cloud IVR is a stepping stone; plenty of operations should simply run a well-built menu in the cloud and stop there. But the migration is when the question is cheapest to answer, and the latency budget you inherit is worth understanding first — the latency calculator shows where the time actually goes, and it is rarely where people assume.