OpenAI GPT-Live API: capabilities, limits, and alternatives

The Rime Team
The Rime Team
September 28, 2026

Five cents a minute. That's what OpenAI charges for the voice layer of GPT-Live-1, the full-duplex voice model it put into its API on September 10, 2026. Full duplex means the model listens and speaks at the same time, so callers can interrupt, laugh, or change their mind mid-sentence and the agent keeps up.

It's the most capable hosted voice agent model OpenAI has shipped, and for plenty of teams it will be the right call. For teams that need to run their voice agent on their own infrastructure, tune it on their own calls, or check every word before a caller hears it, it's the wrong shape. This guide covers what GPT-Live-1 does, what it costs, what changed on compliance, and where the real limits sit.

What GPT-Live-1 is

GPT-Live-1 splits a voice agent into two parts. In OpenAI's words, "GPT-Live handles conversation. It listens, speaks, and decides when to ask the backend for help." Reasoning and tool calls go to a separate backend model, such as GPT-6 Astra or Luna, which you pick independently of the voice model.

There are two ways to wire the backend:

  • Responses Delegation: OpenAI runs the backend model and passes context and results between it and GPT-Live. You supply the function tools.
  • Client Delegation: you connect your own agent and control the backend's context, execution, and results yourself.

Sessions connect over WebRTC for browsers, WebSockets for servers, or SIP for phone calls, and a sideband WebSocket lets you monitor a live session. OpenAI launched it with 14 voices across accents and dialects. Custom voices exist but are limited to eligible customers through OpenAI's sales team.

OpenAI's launch numbers: a 30-point improvement over gpt-realtime-2.1 on Full Duplex Bench, and first place on Tau3, a voice agent intelligence benchmark. Speak, one of the launch customers, reported cutting interruptions by almost 80% compared with turn-based systems. Yelp, Fin, Cognition, Picsart, and HeyGen are also named launch customers. All of these figures come from OpenAI's announcement, and none have been independently reproduced yet.

How GPT-Live relates to the Realtime API

Timeline of OpenAI's Realtime API, from the October 2024 public beta to general availability in August 2025 and the 2026 model releases

GPT-Live sits on top of two years of work on the Realtime API, which is still available and is where most existing OpenAI voice agents run.

The Realtime API launched as a public beta on October 1, 2024, running gpt-4o-realtime-preview over WebSockets. It reached general availability on August 28, 2025 with the gpt-realtime model, SIP phone support, and native remote MCP (Model Context Protocol) support, so a session can point at a remote tool server and use its tools directly. OpenAI removed the beta interface on May 12, 2026, and teams still on it had to migrate.

The 2026 releases moved quickly. OpenAI shipped gpt-realtime-2 on May 7, alongside GPT-Realtime-Translate for streaming speech translation and GPT-Realtime-Whisper for streaming speech-to-text. gpt-realtime-2.1 and a cheaper gpt-realtime-2.1-mini followed on July 6.

The capability jumps along the way:

  • Context grew from about 32,000 tokens on the original gpt-realtime to 128,000 on gpt-realtime-2 and 2.1, with up to 32,000 output tokens.
  • Function calling on the ComplexFuncBench audio eval rose from 49.7% on the December 2024 preview to 66.5% at general availability.
  • OpenAI reports p95 latency down at least 25% across its Realtime voice models, attributed to caching.
  • On Big Bench Audio, a reasoning benchmark, OpenAI reported 82.8% for gpt-realtime at launch. Artificial Analysis later measured gpt-realtime-2 at 96.6%. Those two numbers come from different evaluators, so treat the gap as directional.

GPT-Live and Realtime API pricing

GPT-Live-1 bills the voice layer by the second, at $0.05 a minute. The backend model is billed separately at its own token rates, so the real per-minute cost depends on which backend you choose and how much it thinks.

The Realtime API bills by token. gpt-realtime-2.1 costs $32 per million audio input tokens, $0.40 per million cached input tokens, and $64 per million audio output tokens. The mini model costs $10, $0.30, and $20 for the same three.

Bar comparison of gpt-realtime-2.1 input pricing: $32 per million uncached audio input tokens versus $0.40 per million cached input tokens, an 80x difference

Look at the two input prices again: $32 versus $0.40, an 80x gap. A voice agent that re-sends its system prompt and conversation history uncached on every turn burns money fast, and that difference decides whether a pilot survives its first invoice. With caching wired correctly, the development shop Fora Soft estimates typical model cost at $0.06 to $0.10 per conversational minute.

The mini model fits FAQ handling and simple reservations. The full model earns its price on multi-step banking questions, clinical triage, and anything where a wrong inference has consequences.

On build versus buy, even single sources disagree. Fora Soft puts the point where a self-managed stack beats a managed platform at roughly 10,000 minutes a month in one analysis, and at 50,000 to 150,000 in another. The honest answer depends on your team's engineering capacity as much as your volume.

HIPAA and compliance: what changed

For most of the Realtime API's life, healthcare teams couldn't send patient audio through it under a Business Associate Agreement (BAA), the contract HIPAA requires with any vendor that handles protected health information (PHI). That has changed. OpenAI's HIPAA eligibility page now lists both the Realtime endpoint and GPT-Live sessions as BAA-eligible.

Two details decide whether a deployment actually qualifies. BAA eligibility requires Modified Retention, and standard API endpoints keep data for 30 days by default, so a team that skips the retention setup is out of compliance on day one. Teams using OpenAI models through Azure should confirm Azure's eligibility list separately, since the two companies publish their own.

HIPAA isn't the only rule in play:

  • TCPA: the FCC's February 2024 declaratory ruling treats AI-generated voices as artificial under the Telephone Consumer Protection Act, so outbound calls need prior express consent, and written consent for marketing. Statutory damages run $500 per call, up to $1,500 if willful.
  • NO FAKES Act: the bill targeting unauthorized AI voice replicas cleared the Senate Judiciary Committee on June 18, 2026 and awaits a full Senate vote.
  • PCI DSS: if card numbers pass through a voice session, the payment card rules apply to the whole path.

The three limits that still apply

With HIPAA covered, three limits remain. They're independent, so a team can hit any one without the others.

1. It runs in OpenAI's cloud. GPT-Live and the Realtime API are hosted services with no self-hosted or on-premise option. Defense, government, and many banks require voice data to stay inside their own data center, which rules both out regardless of BAA coverage.

2. The voice layer is OpenAI's. Client Delegation lets you bring your own backend agent, which helps. The part that listens and speaks stays OpenAI's model, though, so a procurement team that requires the option to swap providers can't get it at the voice layer.

3. You can monitor it, but you can't shape it. The sideband connection lets you watch a live session. OpenAI's GPT-Live guide doesn't document a way to review and block a reply before it's spoken, and there's no option to fine-tune the voice model on your own calls. For a bank, where a spoken promise about a rate carries the same weight as a written one, that gap matters.

The usual workaround is to go back to a chained pipeline: a speech-to-text model, a text LLM, and a text-to-speech model, each under its own contract and swappable, orchestrated with a framework like Pipecat or LiveKit Agents. That restores control and costs you what made GPT-Live attractive in the first place. Every handoff adds latency, and the moment speech becomes a transcript, the tone and hesitation in the caller's voice are gone. We cover that tradeoff in what is speech-to-speech?

So today, regulated teams pick between hosted quality and their own control. A speech-to-speech model you can inspect, fine-tune on your own calls, and run on your own infrastructure removes that choice. Rime is building one. More soon.

Where GPT-Live fits by industry

Healthcare. Appointment scheduling, benefits verification, and patient support lines are now possible on OpenAI's voice stack under a BAA, with retention configured correctly. Audio recordings count as PHI, so every vendor in the stack needs a signed BAA, including the telephony provider.

Banking and financial services. Card activation, balance inquiries, lending eligibility, and insurance servicing all need guardrails on what the agent says about rates and eligibility, consent rules that vary by state, and a replayable record of every call. On-premise requirements are common, which is where limit one bites.

Food ordering. This is GPT-Live's cleanest fit, since PHI rarely comes up and the deciding factors are speed and how natural the voice sounds. The category is still early: only about 6% of restaurant operators use AI to take orders, according to the National Restaurant Association's 2026 industry report. Yum! Brands has run more than two million drive-thru orders through voice AI at 300+ Taco Bell locations. Wendy's FreshAI, built with Google Cloud, cut service times by about 22 seconds versus the regional average in its Columbus pilot. Burger King is piloting Patty, an OpenAI-powered headset assistant for staff, in 500 US restaurants. The category also has its cautionary tale: McDonald's ended its IBM automated order-taking test in 2024 after accuracy problems. SoundHound launched OASYS, a platform for orchestrating agents across phones, kiosks, and cars, in May 2026. For outbound calls, TCPA consent is the rule to watch.

Hospitality. Hotels use voice agents for reservations, booking confirmations, upsells, in-stay requests, and overnight front-desk overflow. HIPAA rarely applies. PCI DSS does, the moment a guest reads out a card number.

GPT-Live alternatives

Managed voice agent platforms. Vapi and Retell get teams to production fastest at lower volumes. They don't solve the on-premise requirement.

Model-agnostic orchestration. Pipecat and LiveKit Agents let you swap speech-to-text, LLM, and text-to-speech providers without rewriting the audio layer. This is the standard route for HIPAA-eligible chained pipelines and for procurement teams that require multi-vendor options.

On-premise platforms. Rasa Voice deploys on-premise or in a private cloud and supports touch-tone (DTMF) input for PINs and account numbers. Cognigy, now part of NICE, also offers self-hosted deployment. Rime delivers under 100 milliseconds to first audio in production, offers 500+ voices, is HIPAA and SOC 2 compliant, and deploys in the cloud, in a VPC, or fully on-premise.

Self-hosted open-weight models. Open-weight LLMs such as Llama and Qwen, run inside a chained pipeline, give full deployment control at some cost in reasoning performance.

Should you build on GPT-Live?

If a hosted voice layer at five cents a minute clears your security, compliance, and procurement bar, try it. If any of the three limits apply, you're being asked to trade the quality of your calls for control over them. Don't accept that trade as permanent.