# Ultrafast (OpenAI)

> Source: https://aiwiki.ai/wiki/ultrafast
> Updated: 2026-10-01
> Fact-checked: 2026-10-01
> Categories: AI Inference, Developer Tools, Large Language Models, OpenAI
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Ultrafast (OpenAI)." aiwiki.ai, 1 Oct 2026. https://aiwiki.ai/wiki/ultrafast
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

Ultrafast is [OpenAI](https://aiwiki.ai/wiki/openai)'s premium speed tier, a processing option that generates output tokens faster than the company's Standard, Fast, Flex and Batch tiers at a correspondingly higher price. It exists in two places: as the `ultrafast` value of the `service_tier` parameter in the [OpenAI API](https://aiwiki.ai/wiki/openai_api), and as a speed mode in [Codex](https://aiwiki.ai/wiki/openai_codex) and [ChatGPT Work](https://aiwiki.ai/wiki/chatgpt_work) on specific subscription plans. OpenAI's own documentation calls it "the fastest service tier in the OpenAI API" and tells developers to "use it when speed justifies the higher cost."[1]

The name first appeared on August 13, 2026, when OpenAI announced a limited preview of an Ultrafast tier for [GPT-5.6 Sol](https://aiwiki.ai/wiki/gpt_5_6) running on [Cerebras](https://aiwiki.ai/wiki/cerebras) hardware at up to 750 output tokens per second.[6] Ultrafast became a shipping product on September 29, 2026 at DevDay 2026, where OpenAI made it generally available for [GPT-6 Astra](https://aiwiki.ai/wiki/gpt_6_astra). The DevDay recap describes Ultrafast as "our premium speed tier for workloads where speed matters most" and puts the gain at up to 8x faster token generation (300 tokens per second) in Codex and up to 6x in the API.[2] On the same day OpenAI said an Ultrafast option for [GPT-6.1 Sol](https://aiwiki.ai/wiki/gpt_6_1_sol) was coming soon.[2][7]

As of October 1, 2026, GPT-6 Astra is the only model with published Ultrafast prices, at $60 per million input tokens and $300 per million output tokens for prompts of 272,000 input tokens or fewer, exactly six times Astra's Standard rates on every published line item.[3] The tier is also resold through [Amazon Bedrock](https://aiwiki.ai/wiki/amazon_bedrock), which added it on September 30, 2026.[12]

## Where Ultrafast sits among OpenAI's service tiers

OpenAI's API pricing page groups rates for the flagship models under five tabs: Standard, Batch, Flex, Fast and Ultrafast. Fast mode is the tier renamed from Priority processing on July 30, 2026; either `service_tier: "priority"` or `service_tier: "fast"` selects it.[3][8] Ultrafast is the top rung. GPT-6 Astra's published rates per million tokens on short-context prompts (272,000 input tokens or fewer) as of October 1, 2026 were:[3]

| Tier | Input | Cached input | Cache writes | Output |
| --- | --- | --- | --- | --- |
| Batch | $5.00 | $0.50 | $6.25 | $25.00 |
| Flex | $5.00 | $0.50 | $6.25 | $25.00 |
| Standard | $10.00 | $1.00 | $12.50 | $50.00 |
| Fast | $20.00 | $2.00 | $25.00 | $100.00 |
| Ultrafast | $60.00 | $6.00 | $75.00 | $300.00 |

For prompts above 272,000 input tokens the Ultrafast rates become $120.00 input, $12.00 cached input, $150.00 cache writes and $450.00 output per million tokens, against $20.00, $2.00, $25.00 and $75.00 on Standard. The six-times multiple holds across the long-context rates as well.[3] GPT-6 Astra's model page notes separately that prompts over 272,000 input tokens are "priced at 2x input and cache rates and 1.5x output for the full request," and that cache writes bill at 1.25x the uncached input rate.[9]

The Ultrafast pricing table listed only `gpt-6-astra` as of October 1, 2026. No Ultrafast rate for GPT-6.1 Sol, GPT-6 Sol, GPT-6 Luna or GPT-5.6 Sol had been published.[3]

## Throughput claims

Every speed figure in circulation originates with OpenAI, and the figures are not interchangeable across surfaces. The table below keeps each claim attached to the source that made it and the surface it describes.

| Claim | Model | Surface | Source and date |
| --- | --- | --- | --- |
| Up to 8x faster token generation, 300 tokens per second | GPT-6 Astra (implied by context) | Codex | DevDay 2026 recap, September 29, 2026[2] |
| Up to 6x faster | GPT-6 Astra (implied by context) | OpenAI API | DevDay 2026 recap, September 29, 2026[2] |
| Up to 8x faster token generation than Standard mode | GPT-6 Astra | Codex | ChatGPT and Codex "Speed" documentation, accessed October 1, 2026[5] |
| Up to 8x faster speeds than Standard mode | GPT-6 Astra | OpenAI API | Ultrafast mode guide page description, accessed October 1, 2026[1] |
| Up to 6x faster than Standard | GPT-6 Astra | OpenAI API via Bedrock | Amazon Bedrock model card, accessed October 1, 2026, attributed to OpenAI[13] |
| Up to 6x faster inference in the API, with up to 300 tokens per second | GPT-6 Astra | Amazon Bedrock | AWS What's New post, September 30, 2026, attributed to OpenAI[12] |
| Up to 8x faster token generation compared to its standard speed in Codex | GPT-6.1 Sol | Codex (announced, not shipped) | GPT-6.1 Sol launch post, September 29, 2026[7] |
| Up to 14x faster than Standard processing, up to 750 output tokens per second | GPT-5.6 Sol | OpenAI API (limited preview) | Previewing Ultrafast post, August 13, 2026[6] |

OpenAI's official account posted the DevDay figures on September 29, 2026: "Our premium speed tier, Ultrafast offers up to 8x faster token generation (300 tokens per second) in Codex and up to 6x in the API."[18]

The Codex documentation attaches an explicit caveat to its own multiple: "This comparison measures token generation speed, not billing rates or overall task completion time."[5] The Ultrafast guide's own page description says "Our fastest API service tier, with up to 8x faster speeds than Standard mode," and the changelog's August 13, 2026 entry says the GPT-5.6 Sol preview "runs up to 14x faster than Standard processing," but neither document puts a tokens-per-second figure on GPT-6 Astra Ultrafast, describing the benefit only as reducing "the time between generated output tokens."[1][4]

## Model availability as of 1 October 2026

| Model | OpenAI API | Codex and ChatGPT Work | Published Ultrafast price |
| --- | --- | --- | --- |
| GPT-6 Astra | Broadly available, "available to all API users at low rate limits"[1] | Available on Pro $500 and eligible Enterprise and Edu plans[5] | Yes[3] |
| GPT-5.6 Sol | Preview access only, arranged through an OpenAI account team[1] | Not documented as an Ultrafast model; the `ultrafast` service tier was removed from `gpt-5.6-sol` in the Codex client in a change merged September 21, 2026[19] | No[3] |
| GPT-6.1 Sol | Announced as coming soon on September 29, 2026[2] | Documentation says the model "supports Standard and Fast where available" and that "Ultrafast support for GPT-6.1 Sol is coming later"[5][11] | No[3] |

OpenAI's DevDay recap wording was: "GPT-6 Astra Ultrafast is available today in the API and in ChatGPT Work and Codex on Pro 500 and Enterprise plans. GPT-6.1 Sol Ultrafast is coming soon."[2] The ChatGPT and Codex documentation is more specific about plans, naming "Pro $500 and eligible Enterprise and Edu plans," adding that eligible Enterprise workspaces must be on credit-based or USD usage-based agreements, that "legacy Enterprise plans that rely on rate limits instead of usage-based billing aren't supported," and that "other self-serve plans don't have access to Ultrafast at launch, even with purchased credits."[5]

Pro $500 is the subscription tier OpenAI introduced at the same DevDay. The recap says it "offers our highest usage allowance at 25 times the ChatGPT Plus allowance and includes access to Ultrafast."[2] The Codex pricing page lists Pro plans "at $100, $200, or $500 USD per month" and puts Astra Ultrafast access on the $500 rung only.[10]

## Selecting and billing Ultrafast

In the API, Ultrafast is a per-request choice rather than a plan. A developer sets `model` to `gpt-6-astra` and `service_tier` to `ultrafast` on a Responses API call, over HTTP or on a WebSocket connection; the OpenAI SDKs expose it for JavaScript, Python, Go, Java, Ruby and curl.[1] The September 29, 2026 changelog entry reads: "Added Ultrafast mode for GPT-6 Astra in the Responses API. Use `gpt-6-astra` with `service_tier: \"ultrafast\"` to reduce the time between generated output tokens."[4]

OpenAI recommends a persistent connection, "especially for agentic applications that make many tool calls in quick succession," and warns that "without a persistent connection, network overhead can reduce the latency gains."[1]

In Codex and ChatGPT Work, Ultrafast is a speed mode rather than a token price, and it consumes subscription allowance at a multiplier. OpenAI publishes two different multipliers for the same tier:[5][10]

| Speed mode | Included subscription usage | Purchased credits and Enterprise pay-as-you-go |
| --- | --- | --- |
| Fast | 2.5x | 2x |
| GPT-6 Astra Ultrafast | 8x | 6x |

OpenAI adds that "these billing multipliers don't describe speed increases."[5] On Pro $500, "Ultrafast uses your included usage first, then your available credits after that allowance runs out." For Enterprise workspaces, Ultrafast is off by default and workspace owners enable it for selected users or the whole workspace through workspace permissions; existing per-user spend controls apply.[5] ChatGPT Work and Codex share one pool of pricing, credits and usage limits, and Codex driven by an API key bills at API token prices instead, with the ChatGPT credit multipliers not applying.[5] For reference, Codex's Standard credit rate for GPT-6 Astra is 250 credits per million input tokens, 25 per million cached input tokens and 1,250 per million output tokens.[10]

## Documented limits

Ultrafast ships with lower rate limits than Standard processing at every usage tier above the first, where the two are equal. OpenAI's guide says the tier "is currently available to all API users at low rate limits" and directs organizations with an OpenAI account team to that team for higher limits or for GPT-5.6 Sol preview access. Default Ultrafast token rate limits for GPT-6 Astra are:[1]

| API usage tier | Ultrafast tokens per minute | Standard tokens per minute for comparison[9] |
| --- | --- | --- |
| Tiers 1-3 | 500,000 | 500,000 (tier 1), 1,000,000 (tier 2), 2,000,000 (tier 3) |
| Tier 4 | 1,000,000 | 4,000,000 |
| Tier 5 | 5,000,000 | 40,000,000 |

Data residency is the other hard limit. The guide states: "Ultrafast supports US data residency and global processing only. It does not support EU or other non-US regional processing endpoints."[1] The changelog repeats that it is available "with global processing and US data residency" and that "EU and other regional inference residency aren't supported."[4] On the ChatGPT side, OpenAI says Ultrafast "isn't available to workspaces that require inference residency outside the United States," while noting that "a workspace's location alone doesn't determine eligibility."[5]

Nothing in OpenAI's Ultrafast documentation describes a different context window, a reduced feature set or different model output; the Ultrafast guide covers configuration, pricing and availability only, and GPT-6 Astra's model page gives one context window of 1,050,000 tokens with a 128,000-token output maximum and an April 30, 2026 knowledge cutoff regardless of tier.[1][9]

## On Amazon Bedrock

AWS announced Ultrafast support for GPT-6 Astra on [Amazon Bedrock](https://aiwiki.ai/wiki/amazon_bedrock) on September 30, 2026, a day after OpenAI's launch, saying that "according to OpenAI, Ultrafast delivers up to 6x faster inference in the API, with up to 300 tokens per second," and suggesting it for "real-time coding assistants, interactive agents, and customer-facing experiences."[12]

Bedrock's model card sets Ultrafast prices at six times the corresponding Standard prices and applies the same `"service_tier": "ultrafast"` switch on the Responses API. Global cross-Region inference matches OpenAI's own rates at $60 input, $75 cache write, $6 cache read and $300 output per million tokens on short context; in-Region and US geographic cross-Region routes carry a 10 percent premium, at $66, $82.50, $6.60 and $330. Routing is restricted: on the `bedrock-mantle` endpoint Ultrafast works only in us-east-1, and on `bedrock-runtime` only through the US geographic or global cross-Region inference profiles. The card notes that "Ultrafast does not support regional Mantle access in us-west-2."[13]

## The August 2026 GPT-5.6 Sol preview

The first Ultrafast was a different thing in all but name. On August 13, 2026 OpenAI published "Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed," announcing "an early look at Ultrafast, a new service tier" that ran GPT-5.6 Sol up to 14 times faster than Standard processing and was "launching first in the OpenAI API." That post is the only OpenAI material that names the hardware: "Powered by Cerebras, Ultrafast generates up to 750 output tokens per second." A later section headed "Powered by Cerebras" describes the tier as "the next step in our partnership with Cerebras to bring ultra-low-latency inference to OpenAI's platform."[6]

OpenAI framed the August release as a capacity-constrained experiment. The post says GPT-5.6 Sol on Ultrafast was "available in a limited preview today to a select group of customers" and that OpenAI would "expand access as capacity grows," and lists incident response, fraud and market analysis, customer support, commerce and research iteration as the workloads it was studying. It also reports internal use at OpenAI, including a research workflow that OpenAI says it hopes to tighten from an overnight batch of experiments to "multiple iterations during the workday."[6] The same post carries four customer testimonials, from John Crepezzi (AI Assistants, Jane Street), Courtland Lykins (Product Lead, Voice AI, Podium), Mitch Troyanovsky (co-founder, Basis) and Alex Wang (Applied AI, Rogo).[6]

The preview survives in the API documentation: as of October 1, 2026 the Ultrafast guide still listed "preview access" for GPT-5.6 Sol, routed through OpenAI account teams.[1] On the Codex side it did not: a change to the open-source Codex client titled "Remove the `ultrafast` service tier from `gpt-5.6-sol`" was merged on September 21, 2026, a week before the Astra launch.[19]

## How OpenAI describes the implementation

OpenAI has published no technical explanation of how Ultrafast reaches its throughput for GPT-6 Astra. The API guide and changelog describe the effect ("reduce the time between generated output tokens") and one piece of client-side advice (use WebSockets), and nothing more.[1][4] Cerebras is named only in the August 2026 GPT-5.6 Sol preview post and only for that model.[6] No OpenAI source reviewed here attributes GPT-6 Astra Ultrafast to Cerebras, to custom silicon, to [speculative decoding](https://aiwiki.ai/wiki/speculative_decoding), to [disaggregated serving](https://aiwiki.ai/wiki/disaggregated_serving) or to any other named technique.

MIXED made the same observation about the gap between the marketing figures and the documentation, noting that "neither the changelog nor the guide puts a number on the speed."[16]

## Reception and independent measurement

No independent benchmark of the Ultrafast tier had been published as of October 1, 2026. [Artificial Analysis](https://aiwiki.ai/wiki/artificial_analysis), which measures [tokens per second](https://aiwiki.ai/wiki/tokens_per_second) for commercial endpoints, listed GPT-6 Astra (max) at 51.1 output tokens per second with Standard pricing of $10.00 input and $50.00 output per million tokens when accessed on October 1, 2026, ranking it 143rd of 223 models for speed and calling it "notably slow"; the site carried no Ultrafast endpoint for the model.[14] OpenAI's figures and that measurement are therefore not directly comparable: one is a vendor ceiling on two specific surfaces, the other a third-party measurement of the Standard tier.

VentureBeat's Carl Franzen put the announced 300 tokens per second in market context on launch day, writing that Ultrafast "would sit firmly in the high-speed end of today's model market, but it would not be the outright throughput leader," and citing Artificial Analysis measurements of about 201 tokens per second for Gemini 3.5 Flash, roughly 769 for Mercury 2 and about 1,491 for Celeris-1. His reading was that "the distinction is that OpenAI is offering that 300-token/sec ceiling on its frontier GPT-6-class models, whereas the absolute speed leaders tend to be models optimized specifically for ultra-high-throughput inference." Franzen also derived, and labelled as derived, a hypothetical $12 input and $60 output per million tokens for a GPT-6.1 Sol Ultrafast tier from the announced six-times multiplier, noting that OpenAI's DevDay materials "state the 6X multiplier rather than separately listing those GPT-6.1 Sol Ultrafast token rates."[15]

MIXED's Shane S. Ellison worked through the published table on October 1, 2026 and reached the arithmetic conclusion that "a developer paying $300 per 1M output tokens in the API is buying up to six times the speed for exactly six times the price," and flagged the residency restriction as "a harder limit than it first looks for anyone whose contracts specify where inference happens."[16]

At the keynote itself, Simon Willison's live blog records [Sam Altman](https://aiwiki.ai/wiki/sam_altman) presenting the price as self-evidently worth paying: Ultrafast costs "6x the price of standard" and, in Willison's transcription of Altman, "you know what, it's worth it." Willison also notes that Romain Huet live-coded a ticket-giveaway app on stage using Ultrafast, and that Willison took an earlier 3D-modelling demo to be using it too, and that in the afternoon question-and-answer session Thibault Sottiaux said Ultrafast had "made building impossible things possible," crediting it with helping OpenAI merge the ChatGPT desktop and Codex desktop apps in 28 days.[17]

## References

1. OpenAI. "Ultrafast mode." OpenAI API documentation. Accessed October 1, 2026. https://developers.openai.com/api/docs/guides/ultrafast-mode
2. OpenAI. "DevDay 2026 Recap." September 29, 2026. https://openai.com/index/devday-2026-recap/
3. OpenAI. "Pricing." OpenAI API documentation. Accessed October 1, 2026. https://developers.openai.com/api/docs/pricing
4. OpenAI. "Changelog." OpenAI API documentation. Accessed October 1, 2026. https://developers.openai.com/api/docs/changelog
5. OpenAI. "Speed." ChatGPT and Codex documentation. Accessed October 1, 2026. https://learn.chatgpt.com/docs/agent-configuration/speed
6. OpenAI. "Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed." August 13, 2026. https://openai.com/index/previewing-ultrafast/
7. OpenAI. "Introducing GPT-6.1 Sol." September 29, 2026. https://openai.com/index/introducing-gpt-6-1-sol/
8. OpenAI. "Fast mode." OpenAI API documentation. Accessed October 1, 2026. https://developers.openai.com/api/docs/guides/fast-mode
9. OpenAI. "GPT-6 Astra." OpenAI API documentation. Accessed October 1, 2026. https://developers.openai.com/api/docs/models/gpt-6-astra
10. OpenAI. "Pricing." ChatGPT and Codex documentation. Accessed October 1, 2026. https://learn.chatgpt.com/docs/pricing
11. OpenAI. "Models." ChatGPT and Codex documentation. Accessed October 1, 2026. https://learn.chatgpt.com/docs/models
12. Amazon Web Services. "OpenAI GPT-6 Astra now supports UltraFast mode on Amazon Bedrock." AWS What's New, September 30, 2026. https://aws.amazon.com/about-aws/whats-new/2026/09/openai-gpt-6-astra-ultrafast-on-amazon-bedrock/
13. Amazon Web Services. "GPT-6 Astra." Amazon Bedrock User Guide. Accessed October 1, 2026. https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-openai-gpt-6-astra.html
14. Artificial Analysis. "GPT-6 Astra (max): Intelligence, Performance & Price Analysis." Accessed October 1, 2026. https://artificialanalysis.ai/models/gpt-6-astra
15. Franzen, Carl. "OpenAI's GPT-6.1 Sol offers Astra-like performance at 1/5th price. A new Ultrafast tier clocks at 300 tokens per second." VentureBeat, September 29, 2026. https://venturebeat.com/technology/openais-gpt-6-1-sol-offers-astra-like-performance-at-1-5th-price-a-new-ultrafast-tier-clocks-at-300-tokens-per-second
16. Ellison, Shane S. "GPT-6 Astra's Ultrafast tier costs $300 per million output tokens, six times standard." MIXED, October 1, 2026. https://mixed-news.com/en/gpt-6-astra-ultrafast-tier-300-million-output-tokens/
17. Willison, Simon. "OpenAI DevDay 2026 live blog." September 29, 2026. https://simonwillison.net/2026/Sep/29/openai-devday-2026-live-blog/
18. OpenAI (@OpenAI). "This is Ultrafast." X, September 29, 2026. https://x.com/OpenAI/status/2104993966043320759
19. openai/codex. "Remove the `ultrafast` service tier from `gpt-5.6-sol`." Pull request #47130, merged September 21, 2026. https://github.com/openai/codex/pull/47130

