Speech-to-speech, or S2S, means exactly what it sounds like. Someone talks, a system understands, and a spoken answer comes back. The definition covers only the input and the output, not what happens in between.
That's a wider net than most people assume. Speech technology bridges human communication and AI by combining speech recognition and text-to-speech, and it's those two endpoints, voice in and voice out, that qualify something as "speech-to-speech" in the first place. Nothing in that definition tells you how the system got from one to the other.
The confusion usually starts in one particular place. People hear "speech-to-speech" and picture a specific product or a specific way of building it, like the term names one particular piece of tech. It doesn't. S2S is architecture-agnostic. It describes a shape, not a build.
Which raises a fair question: if the term doesn't tell you how the system works, what good is it? Not much, on its own. Two systems can both wear the "speech-to-speech" label while sharing almost nothing under the hood, no common components, no shared logic, nothing. One might chain together three separate models. The other might run a single model start to finish. Each one takes part in the conversation. Both listen. Both get called S2S.
That gap between the label and the internals is exactly why the architecture running underneath matters more than the definition does. Knowing something is "speech-to-speech" tells you what it does. Knowing which architecture it runs on tells you how it'll fail, where it'll lag, and whether it's ready for a hospital call center or a bank hotline. That's the real story, and it starts with the architecture that runs most of these systems today.
The cascaded pipeline: how most production S2S systems are built today
Picture three workers passing a note down a line, each one translating it into a different language before handing it off. That's roughly how a cascaded voice system works. It's a chain of three separate models, each doing one job, in sequence.
Stage one is speech-to-text, or STT. As the caller talks, the STT model listens and transcribes in real time, spitting out a stream of text word by word. No waiting for the sentence to finish. It works live.
Stage two hands that transcript to a large language model. The LLM reads the text, reasons over it, and generates a reply, also as text. This is the same kind of model behind most text-based AI products, so it can be prompted, given tools to call, and instructed the same way you'd instruct any chatbot. It has no idea it's part of a voice system. As far as it knows, it's reading a transcript someone typed.
Three models, three jobs, one conversation. And this remains a current approach still very much in active use. It's the architecture running most production voice agents right now. Teams build cascaded pipelines on tools like AssemblyAI's Voice Agent API, and platforms including Vapi and LiveKit already run production voice agents on exactly this design.
Why has this setup stuck around? A few reasons, and they're good ones.
Each piece is modular. Want a better speech recognizer? Swap it in. Want a different voice? Swap the TTS model, leave everything else alone. The pipeline also produces something valuable as a byproduct: a text transcript of the whole conversation, useful for captions, records, audits, and search. And when something breaks, you can actually look at the transcript and figure out whether the recognizer misheard something or the LLM reasoned badly. That's a debugging advantage that's hard to overstate.
Solid design, honestly. But it comes with one structural crack running straight through it. TTS converts the LLM's text reply into spoken audio, the final voice the caller hears.
The defining weakness of cascaded systems: error propagation through the chain
If the first runner trips, everyone after them is running the wrong direction, and they don't know it.
That's the cascade's core weakness. If the STT stage mishears a word, that wrong word becomes the LLM's input, and the LLM has no way of knowing the transcript in front of it is already broken. It doesn't second-guess. It just reasons over whatever text it's handed, confidently, as if the transcript were gospel. The mistake doesn't get caught. It gets compounded.
Think about what that means in a real call. A patient says the name of a medication, and the recognizer mishears it. The LLM never hears the audio, only the (wrong) text, so it responds based on a drug that was never mentioned. Downstream, everything treats that wrong transcript as ground truth, because nothing in the chain is built to question it.
There's a second cost too, quieter but constant: latency. Three models means three sequential round-trips, and each one has to fully finish before the next can start⟇⟩. The recognizer has to land on final text before the LLM can start reasoning. The LLM has to finish its reply before TTS can start speaking it. Delay stacks up at every handoff.
And natural conversation doesn't wait politely for a pipeline to catch up. People interrupt each other, trail off, and jump back in mid-thought. Handling that kind of back-and-forth requires a separate turn-taking layer bolted on alongside the LLM, since nothing in the cascade natively understands when a human is about to start or stop talking. It's an add-on, not something built into the core.
None of this is a knock on any particular company's engineering. It's a property of the architecture itself, three separate models means three separate blind spots stacked on top of each other. And it's exactly this weakness that pushed researchers toward a different idea entirely: what if there weren't three models at all?
End-to-end speech-to-speech: one model, audio in, audio out
Strip out the relay race. Replace it with one runner who carries the message the whole way. That's the pitch behind end-to-end speech-to-speech: a single model takes raw audio in and produces audio out, directly, with no text stage sitting in the middle. No transcription step. No separate LLM call. No separate TTS call. One model, doing the whole job.
Industry usage gives it a few different names, native audio models, native S2S models, or realtime models, but they all point at the same idea.
When these models showed up in force in 2024, the reaction in a lot of corners was immediate: the cascade is dead. Why bother chaining three models together when one model can do the whole job, faster and more naturally? It's a fair question on its face. If you can skip the handoffs, you skip the compounding errors and the latency that comes with them.
There's a real advantage buried in that pitch too, and it's not just about speed. Because an end-to-end model holds the actual audio the whole way through, it can, in principle, hang onto things text throws away: tone, emotion, pacing, who's speaking and how they sound. A cascade converts everything to text at stage one, and text has no tone. It's flat by design. End-to-end skips that flattening entirely.
The same logic applies to the messier parts of conversation, backchanneling (the "mm-hm" and "right, go on" that keep a conversation alive) and turn-taking. In an end-to-end model, these come more naturally, because the model is working with continuous audio the whole time, not a separate bolted-on layer trying to guess when someone's about to talk.
So does that settle it? The benchmarking research (Moshi, an open and accessible research prototype, being a good example of the state of things) shows the advantage is narrower than the pitch suggests. It's real on some dimensions. On others, the advantage barely appears in the results at all. The advantage doesn't show up evenly across every dimension.
What the research shows when architectures are measured side by side
This is where a lot of the hype either holds up or falls apart, and the most useful evidence so far comes from a benchmarking framework called COMPASS, published 2 June 2026, the first unified and reproducible framework built specifically to test offline speech-to-speech translation systems side by side.
1,248 model-language configurations, 46 separate metrics, eight evaluation dimensions, ten language pairs, cascaded and end-to-end architectures were all tested under the same conditions. Those eight dimensions covered translation quality, audio naturalness, speaker consistency, prosody and emotion, isochrony (timing), isometry (length matching), and lip sync.
The headline finding cuts against the "one architecture wins" narrative entirely: no single architecture dominates across the board. Cascaded and end-to-end systems show complementary strengths, not a clean hierarchy.
Break it down by dimension and the picture sharpens. On naturalness and speaker preservation, the gap between the best and worst systems exceeds 30%, and end-to-end holds the higher ceiling on both. That tracks with the earlier logic: hold onto the audio, hold onto the voice quality and emotional texture. But on translation quality, the gap between architectures stays within a few points, meaning cascaded pipelines hold their ground where getting the meaning exactly right matters most.
So which one should a team pick? The honest answer from COMPASS is: it depends what you're building. Picking one single metric to crown a winner "systematically misrepresents system quality," according to the study, because the right set of things to measure changes depending on the deployment.
What actually matters shifts by domain. For live interpreting, translation quality is critical and naturalness is secondary. For video dubbing, naturalness, speaker consistency, prosody and emotion, and lip sync are all critical at once. For conversational and medical use, translation quality and naturalness are both critical, the strictest bar for getting the meaning right. For podcasts and audiobooks, naturalness, speaker consistency, and emotional prosody carry the most weight.
Standalone mean opinion score predictors, the common shorthand for "does this sound good to a human," can actually correlate negatively with emotional fidelity, which is a warning against shortcuts. In plain terms, a system can sound pleasant and still miss the emotional truth of what's being said. Measuring voice quality is genuinely harder than it looks.
Put it together and the research doesn't crown a winner. It hands you a map instead, one that says: match the architecture to what you're actually trying to do. Which leads to an uncomfortable question for anyone building a product today. If the map says "it depends," why does almost everyone in production still land in the same spot? Per COMPASS, domain-specific metric subsets reduce the full 46 metrics to 10 per direction while preserving ranking reliability (Spearman's ρ >0.80) and cutting evaluation time by roughly 2.5×.
Why cascaded pipelines still dominate production in 2026 despite end-to-end's promise
Because as of 2026, there's no end-to-end speech-to-speech model widely available as a production API. The research exists. The benchmarks exist. What doesn't exist yet is the boring stuff that actually lets a company bet its customer support line on something: production tooling, uptime guarantees, reliability tested under real load, enterprise support contracts. All of that has been built out for STT APIs, LLM APIs, and TTS APIs separately, over years. It hasn't been built for end-to-end S2S yet.
End-to-end is finding real footing in a narrower lane: short, quick conversational exchanges where speed is everything and the system doesn't need to reach out to five different tools mid-conversation. That's a legitimate use case. It's just not the same use case as a deep enterprise workflow with compliance sign-off attached to it.
There's also a wrinkle in the latency argument that made end-to-end so appealing in the first place. Recent analysis shows optimized cascaded systems can hit competitive latency numbers while keeping the LLM's superior text-based reasoning intact. So the "three models are slower" argument, while true in a vacuum, turns out to be narrower in practice than it first sounded.
Then there's what happens when procurement teams actually kick the tires. Teams evaluating voice AI platforms in 2026 report a familiar pattern: a vendor demo looks sharp, the conversation feels natural, and then the security questionnaire lands, and the platform quietly reveals it wasn't built for regulated industries at all, transcripts logged into a shared cloud tenant, no data processing agreement in place, a SOC 2 Type I certification standing in for the Type II the buyer actually needs.
That gap matters because of the scale involved. Gartner forecasts conversational AI will cut contact center labor costs by $80 billion in 2026, with 80% of businesses planning to bring the technology in. At that scale, reliability is a requirement. It's the whole game.
None of this means end-to-end loses forever. It means that in 2026, if a team is deploying at production scale inside a regulated industry, cascaded is what's actually going out the door. Theory and shipping software are two different things, and right now the gap between them runs straight through the enterprise compliance requirements covered next.
What S2S voice agents do in healthcare, banking, food ordering, and hospitality
Strip away the architecture debate for a second and look at where this stuff is actually running today, because it's already well past the pilot stage in several industries.
In healthcare, voice agents handle appointment management, benefits verification, and patient support lines. Any time that data touches protected health information, HIPAA's Security Rule requirements kick in at every single stage of the pipeline, not just at the end.
In banking, voice agents handle card activation, balance inquiries, and routing for first notice of loss. Every word the system says out loud carries risk, so guardrails have to check outbound statements, rates quoted, promises made, eligibility claims, before they ever get spoken.
Food ordering is where the deployment numbers get genuinely striking. Brands) has processed more than two million drive-thru orders through voice AI across more than 300 locations, with order containment above 90%. Burger King began piloting a voice assistant called Patty across 500 restaurants starting in February 2026, handling order-taking and flagging out-of-stock items to managers in real time. SoundHound's platform powers AI ordering at Chipotle, Applebee's, White Castle, Casey's, Five Guys, Habit Burger, and Firehouse Subs. Presto reports roughly 95% accuracy on drive-thru voice deployments, a 20-second throughput improvement, and about 9 hours a day of labor savings per location.
Hospitality has its own story. Slang AI is deployed at more than 2,000 restaurant locations, handling reservations, private dining, catering questions, and general guest inquiries. It raised $36 million in a Series B round in February 2026, bringing its total funding to $68 million, and markets its product as an AI "Superhost". Across hospitality, travel, and financial services more broadly, well-configured voice deployments are hitting 85 to 90% customer satisfaction on fully resolved calls, with containment rates above 50%.
These are deployments running at scale in production. They're operational, at scale, handling millions of real interactions. Which means the compliance questions are live and current too. They're live, right now, in every one of these deployments. For comparison, the human baseline for order accuracy at a human-run drive-thru averages 80–85%.
Compliance requirements that shape how S2S systems must be built for regulated industries
Healthcare sets the strictest bar of the group. Any voice data touching protected health information triggers HIPAA's Security Rule, which demands administrative, physical, and technical safeguards across the board. Every vendor touching that data chain needs a signed Business Associate Agreement, and audio recordings themselves count as PHI, not just the transcript.
This isn't a theoretical risk. One healthcare provider's voice AI failed its HIPAA audit in 2025 because it retained patient conversations for 90 days when the required window called for deletion, and the fallout was a $2.3 million fine plus a three-week shutdown.
Banking carries its own weight. Every outbound sentence a voice agent speaks, quoted rates, promises, eligibility statements, needs a guardrail checking it before it reaches the caller's ear. Consent and call recording rules aren't uniform either; what's required in one state can differ from the next, so the consent step has to be jurisdiction-aware rather than one-size-fits-all.
A recurring gap appears during enterprise procurement, too. Vendors often carry only a SOC 2 Type I certification, which verifies that controls are designed correctly, not a Type II, which verifies those controls actually work over time under real operating conditions. Add to that an LLM provider sitting in the stack as a third-party subprocessor with no data processing agreement in place, and a buyer's legal team has real reason to pause.
A handful of vendors have built specifically around these requirements. Kore.ai runs a broad enterprise platform with prebuilt templates for banking, insurance, and healthcare, unifying customer service, IT, and HR under one roof. Replicant plugs voice-first deployment directly into existing contact center platforms like Genesys and Five9, positioning itself for insurance carriers, telecoms, and healthcare administrators. Rime runs sub-100 millisecond latency with more than 600 customizable, linguist-engineered voices, supports on-premise and secure-environment deployment, executes Business Associate Agreements, and holds production-grade stability under load, purpose-built for healthcare, banking, food ordering, and hospitality.
That last detail, on-premise deployment, deserves its own close look, because for the strictest buyers, it's a core requirement. It's the whole requirement. Named enterprise voice AI vendors serving regulated industries are carried with details exactly as sourced. Cognigy, Düsseldorf-based and founded in 2016, raised a $100M Series C in 2024 led by Eurazeo, serves Lufthansa, Bosch, Toyota, and European banks and insurers, and handles both voice and chat in a single orchestration layer.
On-premise and secure-environment deployment as the compliance answer for the strictest regulated buyers
Cloud-only platforms run into a hard wall with certain buyers. Most cloud voice platforms store data in a single region or offer only a handful of geographic options, and a shared cloud tenant means the speech recognition model might log transcripts in a place the buyer doesn't fully control. For a hospital system with strict PHI handling rules, or a government agency working with classified data, that's a serious compliance risk. It's disqualifying. These buyers simply cannot use a platform built that way.
So what's the alternative? At the strict end of the spectrum sits fully on-premise deployment: the vendor's software runs entirely inside the buyer's own data center or sovereign cloud, and nothing leaves that perimeter. The voice models, the LLM, and the telephony layer all run on infrastructure the buyer owns and controls. It's the hardest setup to operate and the slowest to stand up. But for the strictest regulatory environments, it's the only version that actually clears the bar.
Somewhere between full cloud and full on-premise sits a hybrid model, blending cloud convenience with on-premise control, with the same security policies expected to hold steady across both environments no matter where a given piece of data happens to sit.
Zoom back out and the throughline across this whole piece becomes clear. Speech-to-speech was never one thing. It's an interface, not an architecture, and the industry's next few years will be spent figuring out, case by case, whether cascaded's transparency or end-to-end's fluency, or some careful blend of both, is the right fit for the job actually on the table.

