Citation and evidence

GPT-Realtime / OpenAI Realtime API

21 min full readUpdated 34 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

OpenAISpeech & Audio AIVoice AI

Cite this article

GPT-Realtime is a family of speech-to-speech models exposed through the OpenAI Realtime API, a low-latency interface that lets developers build voice agents which take spoken audio in and return spoken audio out without an intermediate text transcription stage.[1][2] The API was announced in public beta on October 1, 2024 at the OpenAI DevDay developer conference in San Francisco, alongside the gpt-4o-realtime-preview model.[1][3] After roughly eleven months of iterative beta releases, OpenAI graduated the service to general availability on August 28, 2025 with a new headline model named gpt-realtime, dropping the preview suffix and the gpt-4o- prefix from the model identifier.[2][4] In May 2026 OpenAI shipped a second generation, including gpt-realtime-2, gpt-realtime-translate, and gpt-realtime-whisper, lifting the context window from 32,000 to 128,000 tokens and adding adjustable reasoning effort.[5][6] The Realtime API is the same audio infrastructure that powers Advanced Voice Mode in ChatGPT and forms the model layer used by voice agent platforms such as Vapi, Retell AI, and LiveKit Agents.[7][8] In July 2026 OpenAI released gpt-realtime-2.1 and scheduled the original gpt-realtime family for removal from the API on January 20, 2027, and on September 10, 2026 it made a separate full-duplex voice interface, GPT-Live, generally available in the API alongside the Realtime API rather than in place of it.[13][19][27]

History

Beta launch at DevDay 2024

OpenAI introduced the Realtime API as part of a four-product DevDay slate on October 1, 2024 that also included Prompt Caching, Vision Fine-Tuning, and Model Distillation.[3] The first available model, gpt-4o-realtime-preview-2024-10-01, was a derivative of GPT-4o tuned for low-latency bidirectional audio over a single WebSocket connection.[1] At launch the API supported six voices (alloy, echo, fable, onyx, nova, shimmer) carried over from the existing text-to-speech endpoint and offered function calling, server-side voice activity detection (VAD), and optional input-audio transcription for logging.[1] OpenAI priced the beta at $5 per million text input tokens, $20 per million text output tokens, $100 per million audio input tokens, and $200 per million audio output tokens, which the company estimated at roughly $0.06 per minute of audio input and $0.24 per minute of audio output.[1]

The TechCrunch coverage of DevDay characterised the launch as positioning OpenAI against third-party voice stacks that previously chained a speech-to-text model, a text language model, and a text-to-speech engine.[3] OpenAI demonstrated the API live with Twilio integrations that placed phone calls on stage, and the company stated that third-party voices were not permitted in order to avoid copyright disputes (a reference to the public dispute earlier in 2024 over the "Sky" voice and Scarlett Johansson).[3]

December 2024 snapshot and WebRTC

On December 17, 2024 OpenAI released two new model snapshots, gpt-4o-realtime-preview-2024-12-17 and gpt-4o-mini-realtime-preview-2024-12-17, alongside native WebRTC support.[9][10] The pricing of the full-sized model fell by roughly 60 percent to $40 per million audio input tokens and $80 per million audio output tokens, with $2.50 per million cached audio input tokens.[10] The mini variant was priced at $10 per million input tokens and $20 per million output tokens, ten times cheaper than the original beta.[10] The update extended the maximum session duration from 15 minutes to 30 minutes and added out-of-band concurrent responses, custom input context, and adjustable response timing controls.[9]

The WebRTC option was significant because it eliminated the need for application servers to terminate user audio: a browser or mobile client could now negotiate a peer connection directly with OpenAI using an ephemeral token issued by a developer backend, which OpenAI recommended for any client-facing deployment to keep API keys server-side.[9][11] OpenAI partnered with LiveKit Agents to publish reference architectures that paired LiveKit's WebRTC transport with the Realtime API's WebSocket model interface, a topology that also underpins ChatGPT Advanced Voice Mode.[8]

Expressive voices and intermediate updates

In late October and November 2024 OpenAI added five additional voices, Ash, Ballad, Coral, Sage, and Verse, which the company described as more expressive than the original six and tunable for emotion, accent, and tone.[12] Prompt caching for the Realtime API was extended to support both text and audio cache hits, with text input cache hits priced at a 50 percent discount and audio input cache hits at an 80 percent discount.[12] OpenAI estimated that a typical 15-minute conversation cost about 30 percent less than at the October launch once cache savings were applied.[12]

Through the first half of 2025 the API received a series of smaller revisions. The cap on simultaneous sessions went first, in an update note appended to the launch post: "Update on February 3, 2025: We no longer limit the number of simultaneous sessions on the Realtime API."[1] Two days later OpenAI announced data residency in Europe, reached through a regional endpoint, eu.api.openai.com.[13][33] A further interim snapshot, gpt-4o-realtime-preview-2025-06-03, arrived on June 3, 2025; the changelog lists it as a new model snapshot for gpt-4o-realtime-preview rather than as a feature release.[13]

General availability: gpt-realtime (August 2025)

OpenAI moved the Realtime API out of beta on August 28, 2025 and at the same time introduced a new headline model called simply gpt-realtime (snapshot gpt-realtime-2025-08-28).[2][4] The dropped prefix was deliberate: OpenAI clarified that the new model was not a straight derivative of GPT-4o but a separate speech-to-speech network with its own training data mix.[4] Pricing for gpt-realtime was set at $4 per million text input tokens, $16 per million text output tokens, $32 per million audio input tokens (with $0.40 per million cached), $5 per million image input tokens, and $64 per million audio output tokens, a further roughly 20 percent reduction in audio cost from the December 2024 snapshot.[2][4]

The GA model added native image input, native MCP server tool calling (remote MCP servers can be configured at session creation), and Session Initiation Protocol (SIP) transport for direct integration with telephony providers such as Twilio Elastic SIP Trunking, carrier PBX systems, and desk phones.[2][14] Two new voices, Cedar and Marin, debuted alongside gpt-realtime and were designated by OpenAI as the recommended choices for production deployments.[2][11] On the Big Bench Audio reasoning evaluation the model scored 82.8 percent accuracy, up from 65.6 percent for the December 2024 preview, while instruction following on the MultiChallenge audio benchmark improved correspondingly.[2]

Second generation (May 2026)

On May 7, 2026 OpenAI released three new Realtime API models simultaneously: gpt-realtime-2, gpt-realtime-translate, and gpt-realtime-whisper.[5][13] The flagship gpt-realtime-2 is described by OpenAI as the first voice model with "GPT-5-class reasoning" and exposes five adjustable reasoning effort levels (minimal, low, medium, high, xhigh) that trade latency against quality, with low as the default.[5][6][15] OpenAI's prompting guide for Realtime models documents all five and tells developers to start at low for most production voice agents.[26] The context window expanded from 32,000 to 128,000 tokens, and the model added parallel tool calling with audio "preamble" narration of in-flight tool work (e.g., "let me check that for you").[6]

OpenAI's announcement attributes its two headline gains to different reasoning settings. At high effort, it says, GPT-Realtime-2 "scores 15.2% higher on Big Bench Audio for audio intelligence than GPT-Realtime-1.5"; at xhigh it "scores 13.8% higher on Audio MultiChallenge for instruction following."[5] Artificial Analysis, which runs its own Big Bench Audio evaluation, puts the absolute score at 96.6 percent for the entry it labels GPT-Realtime-2 (High), tied with Gemini 3.1 Flash Live Preview - High.[25] Launch coverage reported the same pairs as 96.6 against 81.4 percent on Big Bench Audio and 48.5 against 34.7 percent on Scale AI's Audio MultiChallenge.[15] Artificial Analysis also publishes a Conversational Dynamics score, a weighted average of pause handling, turn taking, interruption handling, and backchannel handling drawn from Full Duplex Bench v1 and v1.5; there the minimal-reasoning entry reaches 96.1 percent, slightly ahead of the high-reasoning entry at 95.3 percent.[25]

gpt-realtime-translate is purpose-built for real-time speech translation across more than 70 input languages and 13 output languages, priced at $0.034 per minute of audio.[5] gpt-realtime-whisper is described by OpenAI as "a new streaming speech-to-text that transcribes speech live as the speaker talks", priced at $0.017 per minute.[5] The older whisper-1 model was deprecated separately, on August 26, 2026, with removal from the API set for February 26, 2027 and gpt-live-transcribe or gpt-transcribe named as its replacements.[19]

GPT-Realtime-2.1 and the January 2027 retirements (July 2026)

On July 6, 2026 OpenAI released gpt-realtime-2.1 and a distilled gpt-realtime-2.1-mini, describing the update as improving alphanumeric recognition, silence and noise handling, and interruption behavior.[13] The model documentation gives gpt-realtime-2.1 the same shape as gpt-realtime-2: text, audio and image input, text and audio output, a 128,000-token context window, 32,000 maximum output tokens, and a September 30, 2024 knowledge cutoff.[30] It is also the model used in the example session in OpenAI's own Realtime API getting-started guide.[31]

Two weeks later, on July 20, 2026, OpenAI announced the deprecation of the older audio and realtime families. gpt-realtime, gpt-realtime-mini, gpt-4o-realtime and gpt-4o-mini-realtime are scheduled for removal from the API on January 20, 2027, with gpt-realtime-2.1 and gpt-realtime-2.1-mini named as the replacements.[19] The gpt-4o-realtime-preview snapshots had already been shut down on May 7, 2026, and the Realtime API beta interface (the OpenAI-Beta: realtime=v1 header) was removed on May 12, 2026.[19]

Relationship to GPT-Live

On September 10, 2026 OpenAI made gpt-live-1 generally available in the API, so developers now have two OpenAI speech interfaces rather than one.[13] GPT-Live is a full-duplex model that listens and speaks at the same time and delegates reasoning and tool use to a separate backend, either an OpenAI-hosted Responses model or an agent the developer runs; the Realtime API instead keeps speech, reasoning and tool use inside a single model and a single session.[29] The two are billed differently as well: a GPT-Live voice session is priced per minute of session duration and billed per second, with backend model and tool usage charged separately, while a Realtime session is billed per token across text, audio and image modalities.[32]

OpenAI's audio documentation steers new work toward the newer product: "For a new conversational voice application, start with GPT-Live," followed by "Use the Realtime API when you need its session and tool model."[27] A dedicated "Migrate to GPT-Live" guide sets out a Realtime migration path covering the connection procedure, the split between speaking behavior and backend business rules, and moving tool definitions out of session.tools into GPT-Live's delegation configuration.[28]

The Realtime API itself is not retired. OpenAI's deprecations page lists individual Realtime model snapshots and the removed beta interface rather than the API, the Realtime API guide is still part of the documentation set, and OpenAI's voice-agent guide presents GPT-Live, the Realtime API, and a chained speech-to-text plus agent plus text-to-speech pipeline as three architectures to choose between.[19][29][31] OpenAI also warns that the two are not drop-in substitutes at the protocol level: sharing a transport "does not make GPT-Live and Realtime handshakes, credentials, or event formats interchangeable."[27]

Architecture

Transport layers

The Realtime API offers three transports, each suited to a different deployment shape.[11][14]

TransportUse caseRecommended for
WebSocketServer-to-server bidirectional JSON event stream with base64-encoded audio framesBackend services, server-orchestrated agents
WebRTCBrowser/mobile peer connection with built-in jitter buffering and packet-loss concealmentDirect client connections
SIPStandard telephony signalling; OpenAI accepts inbound SIP INVITEs and dispatches realtime.call.incoming webhooksPhone numbers, PBX, IVR replacement

Expanded article table

WebSocket was the only transport at launch; WebRTC was added on December 17, 2024; SIP was added as part of the August 2025 GA release.[9][2] OpenAI documentation recommends WebRTC for any client running outside a controlled data centre because the protocol's loss recovery and adaptive jitter buffering produce more consistent quality on consumer networks than raw WebSocket framing.[11]

Event model

A Realtime session is driven by an event stream defined by approximately three dozen typed events, split into client-emitted events (such as session.update, input_audio_buffer.append, response.create, and conversation.item.create) and server-emitted events (such as session.created, session.updated, input_audio_buffer.speech_started, response.output_audio.delta, response.function_call_arguments.done, and response.done).[16][34] Audio is streamed in 20-millisecond frames as base64-encoded payloads inside JSON events; output audio arrives as a series of response.output_audio.delta events that the client concatenates and plays.[16][34]

Voice activity detection is performed server-side by default. The session config exposes a turn_detection block that lets developers tune the silence threshold and prefix padding; when a user pause crosses the threshold the server emits input_audio_buffer.speech_stopped and automatically triggers a response unless turn detection is disabled.[16][11] As a GA feature OpenAI added an input_audio_buffer.timeout_triggered event that fires after a configurable idle timeout.[11]

Audio formats

Speech-to-speech sessions support three audio formats on both input and output, named in the session configuration as audio/pcm (signed 16-bit little-endian PCM, mono, 24 kHz), audio/pcmu (G.711 mu-law at 8 kHz) and audio/pcma (G.711 A-law at 8 kHz).[34] The separate transcription-session endpoint still names its input formats pcm16, g711_ulaw and g711_alaw.[17] The two G.711 variants are logarithmic codecs used by classical telephony; their inclusion is what allows direct interconnection with SIP trunks and PBX systems without an external transcoder.[34][14] Input and output formats are configured separately, through session.audio.input.format and session.audio.output.format, so a Twilio inbound call can stream audio/pcmu while the application receives 24 kHz audio/pcm for archival.[16][14]

Function and tool calling

Tools are declared at session start in a tools array that mirrors the schema used by the OpenAI API Chat Completions endpoint. When the model decides to call a tool it emits a response.function_call_arguments.delta stream followed by response.function_call_arguments.done; the client executes the tool and writes the result back via a conversation.item.create event with a function_call_output item, then optionally requests a follow-up response.[16] Beginning with the GA release function calling became fully asynchronous: the model can continue conversing while a tool call is in flight, automatically producing filler utterances such as "I'm still waiting on that" rather than blocking on the tool's completion.[11] The Realtime API also reaches remote MCP server endpoints, which are registered as entries in the session's tools array with type: "mcp" and a server_label; their tools then become callable in the conversation.[34]

Voices

The voice catalogue grew across releases. The launch lineup of alloy, echo, fable, onyx, nova, and shimmer came from the pre-existing TTS endpoint.[1] Ash, Ballad, Coral, Sage, and Verse were added in late October 2024 as more expressive options.[12] Cedar and Marin shipped with the August 2025 GA gpt-realtime model and are documented as the recommended voices for assistant audio output.[2][11]

Latency

OpenAI does not publish official end-to-end latency numbers, but third-party measurements during 2025 placed median time-to-first-byte at approximately 500 milliseconds for clients in the contiguous United States, with full-sentence response latencies of roughly 1.2 to 2.0 seconds and 95th-percentile latency creeping to 2.5 to 3.0 seconds under noisy input or long tool chains.[18] LiveKit's published architecture for Advanced Voice Mode reports an end-to-end target of approximately 300 milliseconds for the client-server WebRTC leg.[8]

Models and pricing

The table below summarises the principal Realtime API model snapshots from beta to second generation.

SnapshotReleasedAudio in / out per 1M tokensCached audio inNotes
gpt-4o-realtime-preview-2024-10-012024-10-01$100 / $200n/aPublic beta launch[1]
gpt-4o-realtime-preview-2024-12-172024-12-17$40 / $80$2.50WebRTC; 60% price cut[10]
gpt-4o-mini-realtime-preview-2024-12-172024-12-17$10 / $20$0.30Cheap variant[10]
gpt-4o-realtime-preview-2025-06-032025-06-03$40 / $80$2.50New model snapshot[13]
gpt-realtime (-2025-08-28)2025-08-28$32 / $64$0.40GA; MCP; SIP; image input[2]
gpt-realtime-mini2025-10-06$10 / $20$0.30Cost-efficient variant[13][32]
gpt-realtime-1.52026-02-23$32 / $64$0.40Interim snapshot[13][32]
gpt-realtime-22026-05-07$32 / $64$0.40128K context; reasoning levels[5][6]
gpt-realtime-translate2026-05-07$0.034/minn/aLive translation[5]
gpt-realtime-whisper2026-05-07$0.017/minn/aStreaming STT[5]
gpt-realtime-2.12026-07-06$32 / $64$0.40Alphanumeric, noise and interruption fixes[13][30]
gpt-realtime-2.1-mini2026-07-06$10 / $20$0.30Distilled reasoning variant[13][32]

Expanded article table

Text and image tokens are billed separately from audio. As of September 2026 OpenAI's pricing page lists gpt-realtime-2 and gpt-realtime-2.1 at $4 per million text input tokens, $0.40 cached, and $24 per million text output tokens, plus $5 per million image input tokens; gpt-realtime and gpt-realtime-1.5 keep the older $16 text output rate.[32]

The June 2025 deprecation notice for gpt-4o-realtime-preview-2024-10-01 gave developers a three-month transition window, and that snapshot was shut down on October 10, 2025.[19]

Comparison to competing speech-to-speech APIs

The Realtime API competes most directly with Google's Gemini Live (delivered as part of the Gemini API and refreshed on March 26, 2026 as gemini-3.1-flash-live), xAI's Grok Voice (which adopted the OpenAI Realtime wire protocol to ease migration), Hume AI's Empathic Voice Interface, and Inworld's Realtime API.[20] Both OpenAI and Google operate on a native multimodal architecture in which a single model ingests audio and emits audio without an explicit transcription step, while Hume and Inworld build on top of orchestrated pipelines.[20] Independent benchmarks from Artificial Analysis show gpt-realtime-2 and gemini-3.1-flash-live tied at 96.6 percent on Big Bench Audio at high reasoning settings.[15][25] Pricing comparisons are imprecise because Google bills per minute and OpenAI bills per token, but third-party calculators consistently report Gemini Live as cheaper per minute at comparable reasoning effort.[20]

A second axis of competition runs through orchestration platforms rather than model providers. Vapi and Retell AI offer a higher-level voice agent layer over both OpenAI Realtime and traditional STT-LLM-TTS pipelines, charging a per-minute platform fee on top of underlying model costs.[21] ElevenLabs competes through Conversational AI, which pairs ElevenLabs TTS voices with a configurable backbone LLM and a proprietary turn-taking model.[21] LiveKit Agents is OpenAI's officially partnered open-source framework for building Realtime API applications and is the reference implementation OpenAI itself uses for ChatGPT Advanced Voice Mode.[8]

Use cases

OpenAI's customer documentation and launch posts highlight several deployment patterns.[1][2][22]

Customer support and contact-centre automation is the most-cited application: voice agents handle inbound calls, answer common questions, capture intent before routing to a human, and execute back-office actions through function calls. The combination of SIP transport, async function calling, and image input (for example, a customer holding a damaged product to the phone camera in a hybrid web call) is positioned for this use case.[2][14] OpenAI names Deutsche Telekom as an example, saying the company is building voice support experiences in which customers speak whichever language they find most comfortable "while the model translates the conversation in real time", and that it is testing the model for multilingual voice interactions.[5]

Language tutoring was an early showcase. Speak, a conversational language-learning app, integrated the Realtime API during the public beta to power role-play sessions in which learners practice spoken conversations in a target language; the app was used in the DevDay keynote demonstration.[1] Educational applications use the API to build interactive tutors that explain concepts verbally and adapt pacing to learner response.[22]

Accessibility and assistive technologies use the Realtime API for hands-free interaction, including screen reader replacement, sign-language adjacent voice interfaces, and conversational interfaces for users with motor impairments.[22] Healthify, a nutrition and fitness coaching app, uses the Realtime API to drive its AI coach "Ria" and routes more complex cases to human dietitians.[2]

Other documented deployments include IVR replacement on top of LiveKit Agents and SIP trunking, voice-controlled robotics through WebRTC connections from on-device controllers, and gaming non-player characters that hold open-ended spoken dialogue with players.[8][14]

Reception and criticisms

Coverage of the October 2024 beta launch was broadly positive on capability but skeptical on cost. TechCrunch noted that the per-minute audio pricing was high enough to make most consumer-scale deployments uneconomic until the December 2024 price cuts; the report also called out OpenAI's decision not to require automated disclosure that callers were speaking to an AI, leaving that responsibility to developers.[3] InfoQ's DevDay 2024 coverage emphasised the integration with Twilio as a stage-demo highlight but also raised pricing concerns.[23]

Developer community discussion during the beta period focused on three recurring issues: input transcription accuracy on accented English and non-English speakers, name and proper-noun recognition in tool-call arguments, and unexpectedly high session costs caused by long silence periods being billed as audio input.[12][9] The December 2024 snapshot was credited with improving input reliability but transcription quality remained a community concern through 2025.[9]

After the August 2025 GA release, coverage focused on the production-readiness of the API: the addition of SIP, MCP, and async function calling made the service practical for contact-centre replacement deployments rather than just demos.[2][11] The May 2026 second-generation release attracted notice both for the GPT-5-class reasoning claim and for the open question of how the new "xhigh" reasoning level interacts with the API's latency floor; OpenAI documentation recommends that production deployments default to "low" effort to preserve real-time response latency.[6][26]

Availability

The Realtime API is available to all paying OpenAI API customers globally. OpenAI's data controls page lists /v1/realtime with United States and European (EEA plus Switzerland) storage and processing regions, reached through regional endpoints such as eu.api.openai.com, and names six residency-eligible models: gpt-realtime, gpt-realtime-1.5, gpt-realtime-mini, gpt-realtime-2, gpt-realtime-2.1 and gpt-realtime-2.1-mini. The same page notes that tracing is not currently EU data residency compliant for /v1/realtime.[33] The service is also exposed through Microsoft Azure OpenAI Service in Microsoft Foundry, where the same model snapshots ship with Microsoft's own SLA and regional deployment options.[24] Azure offers separate WebSocket, WebRTC, and SIP transports mirroring the OpenAI public surface.[24]

OpenAI deprecated gpt-4o-realtime-preview-2024-10-01 on June 10, 2025 with a three-month sunset window; the snapshot was shut down on October 10, 2025.[19] As of the May 2026 second-generation release, the gpt-realtime slug pointed to the August 2025 snapshot, the gpt-realtime-2 slug to the May 2026 snapshot, and the gpt-realtime-mini slug to the December 2025 mini snapshot.[13] After the July 2026 deprecation notice, OpenAI's recommended models for new Realtime work are gpt-realtime-2.1 and gpt-realtime-2.1-mini; the gpt-realtime, gpt-realtime-mini, gpt-4o-realtime and gpt-4o-mini-realtime families remain callable until January 20, 2027.[19]

See also

References

  1. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10OpenAI, "Introducing the Realtime API", OpenAI, 2024-10-01. openai.com/...introducing-the-realtime-api. Accessed 2026-05-25.
  2. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14OpenAI, "Introducing gpt-realtime and Realtime API updates for production voice agents", OpenAI, 2025-08-28. openai.com/...introducing-gpt-realtime. Accessed 2026-05-25.
  3. ^1 ^2 ^3 ^4 ^5Kyle Wiggers, "OpenAI's DevDay brings Realtime API and other treats for AI app developers", TechCrunch, 2024-10-01. techcrunch.com/...her-treats-for-ai-app-developers. Accessed 2026-05-25.
  4. ^1 ^2 ^3 ^4Simon Willison, "Introducing gpt-realtime", Simon Willison's Weblog, 2025-09-01. simonwillison.net/...introducing-gpt-realtime. Accessed 2026-05-25.
  5. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10OpenAI, "Advancing voice intelligence with new models in the API", OpenAI, 2026-05-07. openai.com/...elligence-with-new-models-in-the-api. Accessed 2026-05-25.
  6. ^1 ^2 ^3 ^4 ^5DataCamp, "GPT-Realtime-2: A Voice Model with GPT-5-Class Reasoning", DataCamp, 2026-05-08. datacamp.com/...gpt-realtime-2. Accessed 2026-05-25.
  7. ^OpenAI Developers, "Developer notes on the Realtime API", OpenAI Developers Blog, 2025-08-28. developers.openai.com/...realtime-api. Accessed 2026-05-25.
  8. ^1 ^2 ^3 ^4 ^5LiveKit, "OpenAI and LiveKit partner to turn Advanced Voice into an API", LiveKit Blog, 2024-10-01. livekit.com/...nership-advanced-voice-realtime-api. Accessed 2026-05-25.
  9. ^1 ^2 ^3 ^4 ^5 ^6OpenAI Developer Community, "Realtime API updates - WebRTC, cheaper prices, 4o-mini, and more", OpenAI, 2024-12-17. community.openai.com/...1059962. Accessed 2026-05-25.
  10. ^1 ^2 ^3 ^4 ^5OpenAI Developers, "gpt-4o-realtime-preview-2024-12-17 and gpt-4o-mini-realtime-preview-2024-12-17 pricing", OpenAI Developers on X, 2024-12-17. x.com/...1869116963588649326. Accessed 2026-05-25.
  11. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9OpenAI, "Developer notes on the Realtime API", OpenAI Developers, 2025-08-28. developers.openai.com/...realtime-api. Accessed 2026-05-25.
  12. ^1 ^2 ^3 ^4 ^5OpenAI Developer Community, "New Realtime API voices and cache pricing", OpenAI, 2024-10-30. community.openai.com/...998238. Accessed 2026-05-25.
  13. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12OpenAI, "Changelog", OpenAI API Documentation. developers.openai.com/...changelog. Accessed 2026-09-11.
  14. ^1 ^2 ^3 ^4 ^5 ^6Twilio, "Connect the OpenAI Realtime SIP Connector with Twilio Elastic SIP Trunking", Twilio Blog, 2025-08-28. twilio.com/...ai-realtime-api-elastic-sip-trunking. Accessed 2026-05-25.
  15. ^1 ^2 ^3MarkTechPost, "OpenAI Releases Three Realtime Audio Models: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper in the Realtime API", MarkTechPost, 2026-05-08. marktechpost.com/...me-whisper-in-the-realtime-api. Accessed 2026-05-25.
  16. ^1 ^2 ^3 ^4 ^5OpenAI, "Realtime conversations guide", OpenAI API Documentation, 2025-08-28. platform.openai.com/...realtime-conversations. Accessed 2026-05-25.
  17. ^OpenAI, "Realtime transcription session reference", OpenAI API Reference, 2025-08-28. developers.openai.com/...create. Accessed 2026-05-25.
  18. ^Skywork AI, "OpenAI Realtime API Review 2025: Honest Pros and Cons", Skywork AI, 2025-10-15. skywork.ai/...ime-api-review-2025-honest-pros-cons. Accessed 2026-05-25.
  19. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8OpenAI, "Deprecations", OpenAI API Documentation. developers.openai.com/...deprecations. Accessed 2026-09-11.
  20. ^1 ^2 ^3MindStudio, "Real-Time AI Voice Models Compared: GPT Realtime 2, Gemini TTS, Grok, and InWorld", MindStudio, 2026-05-09. mindstudio.ai/...ime-ai-voice-models-compared-2025. Accessed 2026-05-25.
  21. ^1 ^2AssemblyAI, "Best Speech-to-Speech Voice Agent API in 2026", AssemblyAI Blog, 2026-02-15. assemblyai.com/...speech-to-speech-voice-agent-api. Accessed 2026-05-25.
  22. ^1 ^2 ^3Skywork AI, "OpenAI Realtime API Use Cases 2025: 10 Real Examples", Skywork AI, 2025-11-12. skywork.ai/...-api-use-cases-2025-10-real-examples. Accessed 2026-05-25.
  23. ^InfoQ, "OpenAI Developer Day 2024 (SF) Announces Real-Time API, Vision Fine-Tuning, and More", InfoQ, 2024-10-08. infoq.com/...openai-sf-dev-day. Accessed 2026-05-25.
  24. ^1 ^2Microsoft, "Use the GPT Realtime API via SIP", Microsoft Foundry Documentation, 2025-10-15. learn.microsoft.com/...realtime-audio-sip. Accessed 2026-05-25.
  25. ^1 ^2 ^3Artificial Analysis, "Speech to Speech Models and Providers Analysis". artificialanalysis.ai/speech-to-speech. Accessed 2026-09-11.
  26. ^1 ^2OpenAI, "Prompting Realtime models", OpenAI API Documentation. developers.openai.com/...voice-prompting. Accessed 2026-09-11.
  27. ^1 ^2 ^3OpenAI, "Audio and voice", OpenAI API Documentation. developers.openai.com/...audio. Accessed 2026-09-11.
  28. ^OpenAI, "Migrate to GPT-Live", OpenAI API Documentation. developers.openai.com/...live-migration. Accessed 2026-09-11.
  29. ^1 ^2OpenAI, "Voice agents", OpenAI API Documentation. developers.openai.com/...voice-agents. Accessed 2026-09-11.
  30. ^1 ^2OpenAI, "GPT-Realtime-2.1", OpenAI API Documentation. developers.openai.com/...gpt-realtime-2.1. Accessed 2026-09-11.
  31. ^1 ^2OpenAI, "Getting started with the Realtime API", OpenAI API Documentation. developers.openai.com/...realtime. Accessed 2026-09-11.
  32. ^1 ^2 ^3 ^4 ^5OpenAI, "API Pricing", OpenAI API Documentation. developers.openai.com/...pricing. Accessed 2026-09-11.
  33. ^1 ^2OpenAI, "Data controls in the OpenAI platform", OpenAI API Documentation. developers.openai.com/...your-data. Accessed 2026-09-11.
  34. ^1 ^2 ^3 ^4 ^5OpenAI, "Realtime server events", OpenAI API Reference. developers.openai.com/...server-events. Accessed 2026-09-11.

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

6 revisions · v7 · 4,215 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Model snapshots, pricing, benchmark attribution and regional availability checked against OpenAI's documentation and Artificial Analysis's live leaderboard on September 11, 2026.

Cite this page: AI Wiki. "GPT-Realtime / OpenAI Realtime API." aiwiki.ai, updated 11 Sept 2026, fact-checked 11 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/gpt_realtime

Suggest edit