Voice is the least forgiving interface you can build on top of a language model. In a chat window, a response that arrives a second late feels normal. In a spoken conversation, the same delay feels broken. Users talk over the assistant, the assistant talks over them, and the illusion of a natural exchange disappears.

OpenAI recently published an engineering write-up on GPT-Live, the system behind its newer real-time voice experience, and InfoQ's Eran Stiller followed up with an interview with Justin Uberti, OpenAI's head of Realtime AI. Uberti is a long-time WebRTC contributor, which shows in the design choices. The details are useful well beyond voice assistants, because they describe how to keep a latency-critical path fast while the rest of an application does slower, less predictable work.

One rule: protect the live path

The central idea in GPT-Live is a strict separation between two kinds of work. The live path contains only the media pipeline, which moves audio in and out, and the inference loop that produces the model's responses. Everything else runs behind an asynchronous RPC boundary: delegating harder questions to larger frontier models, calling tools, saving conversation state and other application logic.

Uberti summarised the principle as "the voice must flow". The practical effect is that slow or variable operations, such as an external API call or a database write, cannot stall the audio. They complete in the background and feed results back into the conversation when ready.

This sounds simple, but it forces difficult design work. According to Uberti, two components needed special attention: making delegation to frontier models fast enough, and rethinking how voice data reaches OpenAI's safety systems. The benefit of the boundary is that each of these could be optimised in isolation, instead of every change risking a regression in the critical path.

Stateful sessions that can still move

A continuous voice conversation is inherently stateful. The model needs the context of everything said so far, and rebuilding that context on every turn would add latency. OpenAI's answer is dedicated, stateful inference per session: each conversation reserves capacity on a specific model instance.

The obvious downside of pinning sessions to machines is reduced elasticity. You cannot easily drain a server for maintenance or rebalance load if users are glued to it. GPT-Live addresses this by allowing a session's context to migrate to another instance in real time. New sessions are steered to instances with spare capacity, and existing ones can be moved when an instance is being drained or when a conversation approaches its context limit. The result combines the latency benefits of sticky sessions with much of the operational flexibility of stateless services.

Making WebRTC start faster instead of replacing it

On the network side, OpenAI stayed with WebRTC, the standard used by browsers for real-time audio and video. WebRTC is mature, handles packet loss and jitter, and includes a complete media pipeline. Its weakness is connection setup, which traditionally takes several network round trips before audio can flow.

OpenAI introduced a set of handshake improvements it calls WARP, short for WebRTC Abridged Roundtrip Protocol, along with a feature called Instant Connect, both aimed at cutting startup latency. Uberti told InfoQ that WARP consists of pieces named SPED, DTLS 1.3 and SNAP, and that each can be deployed independently, so the team could measure the effect of each change on its own. Secondary coverage reports that the combined changes reduce media and data startup from six round trips to one, and that OpenAI is pursuing standardisation through the IETF, though those details come from reports on the original post rather than the interview itself.

Why not switch to something newer, such as media over QUIC? Uberti's answer is pragmatic. Those efforts currently provide only a transport layer, not a full media stack, and still lack features such as GCC congestion control and round-trip-aware path selection. Improving WebRTC's handshake was judged less risky than replacing the transport on both client and server. A side benefit is that existing WebRTC applications can gain from these improvements without code changes. He also expects WebRTC and QUIC to converge over time.

Testing with real traffic, without users noticing

Perhaps the most transferable lesson is how OpenAI validated the system before launch. Instead of relying only on synthetic load tests with recorded speech, the team ran a "silent" test. Real incoming voice traffic from the existing product was mirrored into GPT-Live, processed by the model, and the output was discarded. The application service ran in an effectively read-only mode, without user credentials, so nothing could leak back to customers.

This exposed problems that synthetic tests had missed. Performance degraded under real load in unexpected ways. One example Uberti gave: in certain regions, some GPUs were not located close to the CPUs feeding them, adding latency that only showed up with authentic, geographically diverse traffic. The fixes could then be validated against the same real traffic before launch day.

What software teams can take from this

Most teams are not building voice assistants, but many build systems where one path must stay fast while other work is slow or unreliable. Checkout flows, trading screens, multiplayer games, collaborative editors and live dashboards all share this shape. Several lessons carry over directly.

  1. Define your critical path explicitly and keep it small. Decide which operations directly affect what the user experiences in real time, and move everything else behind an asynchronous boundary.

  2. Do not let optional work block mandatory work. Tool calls, analytics, persistence and enrichment should never be able to stall the core interaction. Design for results that arrive later.

  3. Stateful does not have to mean stuck. If you need sticky sessions for performance, invest early in the ability to migrate them. It pays off in deployments, scaling and incident response.

  4. Improve proven technology before replacing it. Extending a mature, well-understood stack in small, independently measurable steps is often safer than a rewrite on a new protocol.

  5. Use shadow traffic before launch. Mirroring real production requests into a new system, with outputs discarded and side effects disabled, catches problems that synthetic tests cannot model, such as real user diversity, geography and infrastructure quirks.

  6. Measure the tail, not just the average. In real-time systems, the slowest few percent of interactions define how the product feels.

GPT-Live is a reminder that much of what makes AI products feel good is classic distributed systems engineering. The model matters, but so do boundaries, state management, network handshakes and honest testing against real traffic.


Source: Eran Stiller, InfoQ, original article linked below.