Citation and evidence

Ultrafast (OpenAI)

15 min full readUpdated 19 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI InferenceDeveloper ToolsLarge Language ModelsOpenAI

Cite this article

Ultrafast is OpenAI's premium speed tier, a processing option that generates output tokens faster than the company's Standard, Fast, Flex and Batch tiers at a correspondingly higher price. It exists in two places: as the ultrafast value of the service_tier parameter in the OpenAI API, and as a speed mode in Codex and ChatGPT Work on specific subscription plans. OpenAI's own documentation calls it "the fastest service tier in the OpenAI API" and tells developers to "use it when speed justifies the higher cost."[1]

The name first appeared on August 13, 2026, when OpenAI announced a limited preview of an Ultrafast tier for GPT-5.6 Sol running on Cerebras hardware at up to 750 output tokens per second.[6] Ultrafast became a shipping product on September 29, 2026 at DevDay 2026, where OpenAI made it generally available for GPT-6 Astra. The DevDay recap describes Ultrafast as "our premium speed tier for workloads where speed matters most" and puts the gain at up to 8x faster token generation (300 tokens per second) in Codex and up to 6x in the API.[2] On the same day OpenAI said an Ultrafast option for GPT-6.1 Sol was coming soon.[2][7]

As of October 1, 2026, GPT-6 Astra is the only model with published Ultrafast prices, at $60 per million input tokens and $300 per million output tokens for prompts of 272,000 input tokens or fewer, exactly six times Astra's Standard rates on every published line item.[3] The tier is also resold through Amazon Bedrock, which added it on September 30, 2026.[12]

Where Ultrafast sits among OpenAI's service tiers

OpenAI's API pricing page groups rates for the flagship models under five tabs: Standard, Batch, Flex, Fast and Ultrafast. Fast mode is the tier renamed from Priority processing on July 30, 2026; either service_tier: "priority" or service_tier: "fast" selects it.[3][8] Ultrafast is the top rung. GPT-6 Astra's published rates per million tokens on short-context prompts (272,000 input tokens or fewer) as of October 1, 2026 were:[3]

TierInputCached inputCache writesOutput
Batch$5.00$0.50$6.25$25.00
Flex$5.00$0.50$6.25$25.00
Standard$10.00$1.00$12.50$50.00
Fast$20.00$2.00$25.00$100.00
Ultrafast$60.00$6.00$75.00$300.00

Expanded article table

For prompts above 272,000 input tokens the Ultrafast rates become $120.00 input, $12.00 cached input, $150.00 cache writes and $450.00 output per million tokens, against $20.00, $2.00, $25.00 and $75.00 on Standard. The six-times multiple holds across the long-context rates as well.[3] GPT-6 Astra's model page notes separately that prompts over 272,000 input tokens are "priced at 2x input and cache rates and 1.5x output for the full request," and that cache writes bill at 1.25x the uncached input rate.[9]

The Ultrafast pricing table listed only gpt-6-astra as of October 1, 2026. No Ultrafast rate for GPT-6.1 Sol, GPT-6 Sol, GPT-6 Luna or GPT-5.6 Sol had been published.[3]

Throughput claims

Every speed figure in circulation originates with OpenAI, and the figures are not interchangeable across surfaces. The table below keeps each claim attached to the source that made it and the surface it describes.

ClaimModelSurfaceSource and date
Up to 8x faster token generation, 300 tokens per secondGPT-6 Astra (implied by context)CodexDevDay 2026 recap, September 29, 2026[2]
Up to 6x fasterGPT-6 Astra (implied by context)OpenAI APIDevDay 2026 recap, September 29, 2026[2]
Up to 8x faster token generation than Standard modeGPT-6 AstraCodexChatGPT and Codex "Speed" documentation, accessed October 1, 2026[5]
Up to 8x faster speeds than Standard modeGPT-6 AstraOpenAI APIUltrafast mode guide page description, accessed October 1, 2026[1]
Up to 6x faster than StandardGPT-6 AstraOpenAI API via BedrockAmazon Bedrock model card, accessed October 1, 2026, attributed to OpenAI[13]
Up to 6x faster inference in the API, with up to 300 tokens per secondGPT-6 AstraAmazon BedrockAWS What's New post, September 30, 2026, attributed to OpenAI[12]
Up to 8x faster token generation compared to its standard speed in CodexGPT-6.1 SolCodex (announced, not shipped)GPT-6.1 Sol launch post, September 29, 2026[7]
Up to 14x faster than Standard processing, up to 750 output tokens per secondGPT-5.6 SolOpenAI API (limited preview)Previewing Ultrafast post, August 13, 2026[6]

Expanded article table

OpenAI's official account posted the DevDay figures on September 29, 2026: "Our premium speed tier, Ultrafast offers up to 8x faster token generation (300 tokens per second) in Codex and up to 6x in the API."[18]

The Codex documentation attaches an explicit caveat to its own multiple: "This comparison measures token generation speed, not billing rates or overall task completion time."[5] The Ultrafast guide's own page description says "Our fastest API service tier, with up to 8x faster speeds than Standard mode," and the changelog's August 13, 2026 entry says the GPT-5.6 Sol preview "runs up to 14x faster than Standard processing," but neither document puts a tokens-per-second figure on GPT-6 Astra Ultrafast, describing the benefit only as reducing "the time between generated output tokens."[1][4]

Model availability as of 1 October 2026

ModelOpenAI APICodex and ChatGPT WorkPublished Ultrafast price
GPT-6 AstraBroadly available, "available to all API users at low rate limits"[1]Available on Pro $500 and eligible Enterprise and Edu plans[5]Yes[3]
GPT-5.6 SolPreview access only, arranged through an OpenAI account team[1]Not documented as an Ultrafast model; the ultrafast service tier was removed from gpt-5.6-sol in the Codex client in a change merged September 21, 2026[19]No[3]
GPT-6.1 SolAnnounced as coming soon on September 29, 2026[2]Documentation says the model "supports Standard and Fast where available" and that "Ultrafast support for GPT-6.1 Sol is coming later"[5][11]No[3]

Expanded article table

OpenAI's DevDay recap wording was: "GPT-6 Astra Ultrafast is available today in the API and in ChatGPT Work and Codex on Pro 500 and Enterprise plans. GPT-6.1 Sol Ultrafast is coming soon."[2] The ChatGPT and Codex documentation is more specific about plans, naming "Pro $500 and eligible Enterprise and Edu plans," adding that eligible Enterprise workspaces must be on credit-based or USD usage-based agreements, that "legacy Enterprise plans that rely on rate limits instead of usage-based billing aren't supported," and that "other self-serve plans don't have access to Ultrafast at launch, even with purchased credits."[5]

Pro $500 is the subscription tier OpenAI introduced at the same DevDay. The recap says it "offers our highest usage allowance at 25 times the ChatGPT Plus allowance and includes access to Ultrafast."[2] The Codex pricing page lists Pro plans "at $100, $200, or $500 USD per month" and puts Astra Ultrafast access on the $500 rung only.[10]

Selecting and billing Ultrafast

In the API, Ultrafast is a per-request choice rather than a plan. A developer sets model to gpt-6-astra and service_tier to ultrafast on a Responses API call, over HTTP or on a WebSocket connection; the OpenAI SDKs expose it for JavaScript, Python, Go, Java, Ruby and curl.[1] The September 29, 2026 changelog entry reads: "Added Ultrafast mode for GPT-6 Astra in the Responses API. Use gpt-6-astra with service_tier: \"ultrafast\" to reduce the time between generated output tokens."[4]

OpenAI recommends a persistent connection, "especially for agentic applications that make many tool calls in quick succession," and warns that "without a persistent connection, network overhead can reduce the latency gains."[1]

In Codex and ChatGPT Work, Ultrafast is a speed mode rather than a token price, and it consumes subscription allowance at a multiplier. OpenAI publishes two different multipliers for the same tier:[5][10]

Speed modeIncluded subscription usagePurchased credits and Enterprise pay-as-you-go
Fast2.5x2x
GPT-6 Astra Ultrafast8x6x

Expanded article table

OpenAI adds that "these billing multipliers don't describe speed increases."[5] On Pro $500, "Ultrafast uses your included usage first, then your available credits after that allowance runs out." For Enterprise workspaces, Ultrafast is off by default and workspace owners enable it for selected users or the whole workspace through workspace permissions; existing per-user spend controls apply.[5] ChatGPT Work and Codex share one pool of pricing, credits and usage limits, and Codex driven by an API key bills at API token prices instead, with the ChatGPT credit multipliers not applying.[5] For reference, Codex's Standard credit rate for GPT-6 Astra is 250 credits per million input tokens, 25 per million cached input tokens and 1,250 per million output tokens.[10]

Documented limits

Ultrafast ships with lower rate limits than Standard processing at every usage tier above the first, where the two are equal. OpenAI's guide says the tier "is currently available to all API users at low rate limits" and directs organizations with an OpenAI account team to that team for higher limits or for GPT-5.6 Sol preview access. Default Ultrafast token rate limits for GPT-6 Astra are:[1]

API usage tierUltrafast tokens per minuteStandard tokens per minute for comparison[9]
Tiers 1-3500,000500,000 (tier 1), 1,000,000 (tier 2), 2,000,000 (tier 3)
Tier 41,000,0004,000,000
Tier 55,000,00040,000,000

Expanded article table

Data residency is the other hard limit. The guide states: "Ultrafast supports US data residency and global processing only. It does not support EU or other non-US regional processing endpoints."[1] The changelog repeats that it is available "with global processing and US data residency" and that "EU and other regional inference residency aren't supported."[4] On the ChatGPT side, OpenAI says Ultrafast "isn't available to workspaces that require inference residency outside the United States," while noting that "a workspace's location alone doesn't determine eligibility."[5]

Nothing in OpenAI's Ultrafast documentation describes a different context window, a reduced feature set or different model output; the Ultrafast guide covers configuration, pricing and availability only, and GPT-6 Astra's model page gives one context window of 1,050,000 tokens with a 128,000-token output maximum and an April 30, 2026 knowledge cutoff regardless of tier.[1][9]

On Amazon Bedrock

AWS announced Ultrafast support for GPT-6 Astra on Amazon Bedrock on September 30, 2026, a day after OpenAI's launch, saying that "according to OpenAI, Ultrafast delivers up to 6x faster inference in the API, with up to 300 tokens per second," and suggesting it for "real-time coding assistants, interactive agents, and customer-facing experiences."[12]

Bedrock's model card sets Ultrafast prices at six times the corresponding Standard prices and applies the same "service_tier": "ultrafast" switch on the Responses API. Global cross-Region inference matches OpenAI's own rates at $60 input, $75 cache write, $6 cache read and $300 output per million tokens on short context; in-Region and US geographic cross-Region routes carry a 10 percent premium, at $66, $82.50, $6.60 and $330. Routing is restricted: on the bedrock-mantle endpoint Ultrafast works only in us-east-1, and on bedrock-runtime only through the US geographic or global cross-Region inference profiles. The card notes that "Ultrafast does not support regional Mantle access in us-west-2."[13]

The August 2026 GPT-5.6 Sol preview

The first Ultrafast was a different thing in all but name. On August 13, 2026 OpenAI published "Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed," announcing "an early look at Ultrafast, a new service tier" that ran GPT-5.6 Sol up to 14 times faster than Standard processing and was "launching first in the OpenAI API." That post is the only OpenAI material that names the hardware: "Powered by Cerebras, Ultrafast generates up to 750 output tokens per second." A later section headed "Powered by Cerebras" describes the tier as "the next step in our partnership with Cerebras to bring ultra-low-latency inference to OpenAI's platform."[6]

OpenAI framed the August release as a capacity-constrained experiment. The post says GPT-5.6 Sol on Ultrafast was "available in a limited preview today to a select group of customers" and that OpenAI would "expand access as capacity grows," and lists incident response, fraud and market analysis, customer support, commerce and research iteration as the workloads it was studying. It also reports internal use at OpenAI, including a research workflow that OpenAI says it hopes to tighten from an overnight batch of experiments to "multiple iterations during the workday."[6] The same post carries four customer testimonials, from John Crepezzi (AI Assistants, Jane Street), Courtland Lykins (Product Lead, Voice AI, Podium), Mitch Troyanovsky (co-founder, Basis) and Alex Wang (Applied AI, Rogo).[6]

The preview survives in the API documentation: as of October 1, 2026 the Ultrafast guide still listed "preview access" for GPT-5.6 Sol, routed through OpenAI account teams.[1] On the Codex side it did not: a change to the open-source Codex client titled "Remove the ultrafast service tier from gpt-5.6-sol" was merged on September 21, 2026, a week before the Astra launch.[19]

How OpenAI describes the implementation

OpenAI has published no technical explanation of how Ultrafast reaches its throughput for GPT-6 Astra. The API guide and changelog describe the effect ("reduce the time between generated output tokens") and one piece of client-side advice (use WebSockets), and nothing more.[1][4] Cerebras is named only in the August 2026 GPT-5.6 Sol preview post and only for that model.[6] No OpenAI source reviewed here attributes GPT-6 Astra Ultrafast to Cerebras, to custom silicon, to speculative decoding, to disaggregated serving or to any other named technique.

MIXED made the same observation about the gap between the marketing figures and the documentation, noting that "neither the changelog nor the guide puts a number on the speed."[16]

Reception and independent measurement

No independent benchmark of the Ultrafast tier had been published as of October 1, 2026. Artificial Analysis, which measures tokens per second for commercial endpoints, listed GPT-6 Astra (max) at 51.1 output tokens per second with Standard pricing of $10.00 input and $50.00 output per million tokens when accessed on October 1, 2026, ranking it 143rd of 223 models for speed and calling it "notably slow"; the site carried no Ultrafast endpoint for the model.[14] OpenAI's figures and that measurement are therefore not directly comparable: one is a vendor ceiling on two specific surfaces, the other a third-party measurement of the Standard tier.

VentureBeat's Carl Franzen put the announced 300 tokens per second in market context on launch day, writing that Ultrafast "would sit firmly in the high-speed end of today's model market, but it would not be the outright throughput leader," and citing Artificial Analysis measurements of about 201 tokens per second for Gemini 3.5 Flash, roughly 769 for Mercury 2 and about 1,491 for Celeris-1. His reading was that "the distinction is that OpenAI is offering that 300-token/sec ceiling on its frontier GPT-6-class models, whereas the absolute speed leaders tend to be models optimized specifically for ultra-high-throughput inference." Franzen also derived, and labelled as derived, a hypothetical $12 input and $60 output per million tokens for a GPT-6.1 Sol Ultrafast tier from the announced six-times multiplier, noting that OpenAI's DevDay materials "state the 6X multiplier rather than separately listing those GPT-6.1 Sol Ultrafast token rates."[15]

MIXED's Shane S. Ellison worked through the published table on October 1, 2026 and reached the arithmetic conclusion that "a developer paying $300 per 1M output tokens in the API is buying up to six times the speed for exactly six times the price," and flagged the residency restriction as "a harder limit than it first looks for anyone whose contracts specify where inference happens."[16]

At the keynote itself, Simon Willison's live blog records Sam Altman presenting the price as self-evidently worth paying: Ultrafast costs "6x the price of standard" and, in Willison's transcription of Altman, "you know what, it's worth it." Willison also notes that Romain Huet live-coded a ticket-giveaway app on stage using Ultrafast, and that Willison took an earlier 3D-modelling demo to be using it too, and that in the afternoon question-and-answer session Thibault Sottiaux said Ultrafast had "made building impossible things possible," crediting it with helping OpenAI merge the ChatGPT desktop and Codex desktop apps in 28 days.[17]

References

  1. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12OpenAI. "Ultrafast mode." OpenAI API documentation. Accessed October 1, 2026. developers.openai.com/...ultrafast-mode
  2. ^1 ^2 ^3 ^4 ^5 ^6 ^7OpenAI. "DevDay 2026 Recap." September 29, 2026. openai.com/...devday-2026-recap
  3. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8OpenAI. "Pricing." OpenAI API documentation. Accessed October 1, 2026. developers.openai.com/...pricing
  4. ^1 ^2 ^3 ^4OpenAI. "Changelog." OpenAI API documentation. Accessed October 1, 2026. developers.openai.com/...changelog
  5. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10OpenAI. "Speed." ChatGPT and Codex documentation. Accessed October 1, 2026. learn.chatgpt.com/...speed
  6. ^1 ^2 ^3 ^4 ^5 ^6OpenAI. "Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed." August 13, 2026. openai.com/...previewing-ultrafast
  7. ^1 ^2OpenAI. "Introducing GPT-6.1 Sol." September 29, 2026. openai.com/...introducing-gpt-6-1-sol
  8. ^OpenAI. "Fast mode." OpenAI API documentation. Accessed October 1, 2026. developers.openai.com/...fast-mode
  9. ^1 ^2 ^3OpenAI. "GPT-6 Astra." OpenAI API documentation. Accessed October 1, 2026. developers.openai.com/...gpt-6-astra
  10. ^1 ^2 ^3OpenAI. "Pricing." ChatGPT and Codex documentation. Accessed October 1, 2026. learn.chatgpt.com/...pricing
  11. ^OpenAI. "Models." ChatGPT and Codex documentation. Accessed October 1, 2026. learn.chatgpt.com/...models
  12. ^1 ^2 ^3Amazon Web Services. "OpenAI GPT-6 Astra now supports UltraFast mode on Amazon Bedrock." AWS What's New, September 30, 2026. aws.amazon.com/...stra-ultrafast-on-amazon-bedrock
  13. ^1 ^2Amazon Web Services. "GPT-6 Astra." Amazon Bedrock User Guide. Accessed October 1, 2026. docs.aws.amazon.com/...model-card-openai-gpt-6-astra
  14. ^Artificial Analysis. "GPT-6 Astra (max): Intelligence, Performance & Price Analysis." Accessed October 1, 2026. artificialanalysis.ai/...gpt-6-astra
  15. ^Franzen, Carl. "OpenAI's GPT-6.1 Sol offers Astra-like performance at 1/5th price. A new Ultrafast tier clocks at 300 tokens per second." VentureBeat, September 29, 2026. venturebeat.com/...clocks-at-300-tokens-per-second
  16. ^1 ^2Ellison, Shane S. "GPT-6 Astra's Ultrafast tier costs $300 per million output tokens, six times standard." MIXED, October 1, 2026. mixed-news.com/...t-tier-300-million-output-tokens
  17. ^Willison, Simon. "OpenAI DevDay 2026 live blog." September 29, 2026. simonwillison.net/...openai-devday-2026-live-blog
  18. ^OpenAI (@OpenAI). "This is Ultrafast." X, September 29, 2026. x.com/...2104993966043320759
  19. ^1 ^2openai/codex. "Remove the `ultrafast` service tier from `gpt-5.6-sol`." Pull request #47130, merged September 21, 2026. github.com/...47130

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

1 revision · v2 · 3,035 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independently fact-checked 1 Oct 2026 against OpenAI's Ultrafast guide, API pricing page, changelog, Codex docs, the DevDay recap, AWS Bedrock and Artificial Analysis; 4 defects corrected incl. a false claim that the API guide states no throughput figure

Cite this page: AI Wiki. "Ultrafast (OpenAI)." aiwiki.ai, updated 1 Oct 2026, fact-checked 1 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/ultrafast

Suggest edit