# GPT-5.3

> Source: https://aiwiki.ai/wiki/gpt-5.3
> Updated: 2026-06-27
> Categories: AI Models, Large Language Models, OpenAI
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution to "AI Wiki (aiwiki.ai)".

GPT-5.3 is a family of [large language models](/wiki/large_language_model) released by [OpenAI](/wiki/openai) in February and March 2026 as part of the [GPT-5](/wiki/gpt-5) lineage. It comprises three distinct models that arrived within a five-week window: GPT-5.3-Codex (February 5, 2026), an agentic coding and knowledge-work model; GPT-5.3-Codex-Spark (February 12, 2026), an ultra-low-latency variant that became OpenAI's first model served on non-NVIDIA silicon; and GPT-5.3 Instant (March 3, 2026), the new default conversational model for [ChatGPT](/wiki/chatgpt). [1][3][4] GPT-5.3-Codex was the first OpenAI model that the company described as instrumental in building itself, and the first launch OpenAI treated as "High" capability for cybersecurity under its Preparedness Framework. [1][2]

## Overview

GPT-5.3 is a family of large language models released by OpenAI across February and March 2026. The three models arrived in quick succession: GPT-5.3-Codex on February 5, GPT-5.3-Codex-Spark on February 12, and GPT-5.3 Instant on March 3. [1][3][4] The release cycle did two things at once. It unified the general and coding branches of the GPT-5 lineage, and it pushed the conversational default toward a less moralizing tone after months of user complaints about [GPT-5.2](/wiki/gpt-5.2) Instant. [4][13] GPT-5.3-Codex was also the first OpenAI model classified as High capability for cybersecurity under the company's Preparedness Framework, which triggered an expanded safety stack at launch. [1][2][12]

The family was current for roughly two months. GPT-5.3 Instant was replaced as the default ChatGPT model by [GPT-5.5](/wiki/gpt-5.5) Instant on May 5, 2026, although paid subscribers retained access to GPT-5.3 Instant for a three-month transition period. [27]

## What is GPT-5.3?

The GPT-5 series began in August 2025 and expanded rapidly through the first half of 2026. [GPT-5.1](/wiki/gpt-5.1) (November 2025) improved instruction following and added more granular personality settings to ChatGPT. GPT-5.2, released in December 2025, introduced GPT-5.2 Instant, GPT-5.2 Thinking, and GPT-5.2 Pro as separate offerings tuned for speed, reasoning, and professional work. GPT-5.2-Codex, a dedicated coding agent, followed on January 14, 2026. [7]

By early 2026 OpenAI was running two parallel tracks. The conversational track (the Instant series) handled everyday ChatGPT traffic. The coding track (the Codex series) handled long-running software engineering tasks through the Codex CLI, the Codex web app, and Codex API endpoints. Each had its own architecture optimizations, pricing, and rollout schedule. The GPT-5.3 release window collapsed that distinction for the Codex line. OpenAI trained a single model that combined the frontier coding performance of GPT-5.2-Codex with the general reasoning and professional knowledge of GPT-5.2. [1] It was also the point at which the conversational Instant model received a high-profile behavioral overhaul rather than a pure capability upgrade. [4]

The sequencing matters because the three releases told three different stories. GPT-5.3-Codex was a capability story, with new state-of-the-art numbers on agentic coding benchmarks. [1][9] GPT-5.3-Codex-Spark was a hardware story, the first OpenAI model running on non-NVIDIA silicon. [3][24] GPT-5.3 Instant was a tone story, the first time OpenAI publicly characterized a model update as primarily about reducing how annoying ChatGPT felt to talk to. [4][13]

## When was GPT-5.3 released?

All three GPT-5.3 models launched within a five-week window in early 2026. [1][3][4]

| Model | Release date | Primary focus |
|---|---|---|
| GPT-5.3-Codex | February 5, 2026 | General agentic coding and knowledge work |
| GPT-5.3-Codex-Spark | February 12, 2026 | Real-time inline coding, ultra-low latency |
| GPT-5.3 Instant | March 3, 2026 | Everyday conversational tasks, default ChatGPT model |

GPT-5.3-Codex launched at a moment of intense competition. The Anthropic Claude Opus 4.6 release fell within the same week, which meant developer publications spent most of February running comparative reviews. [12] GPT-5.3 Instant became the default model for all ChatGPT tiers on its release date, replacing GPT-5.2 Instant. GPT-5.2 Instant remained available to paid subscribers in a three-month transition period intended to ease workflow disruption for power users who had grown attached to its tone. [4]

## What are the GPT-5.3 variants?

### GPT-5.3-Codex

GPT-5.3-Codex was OpenAI's most capable agentic coding model at the time of its release. [1] Unlike its predecessor, which was a specialist fine-tune on the coding domain, GPT-5.3-Codex was trained jointly on the GPT-5.2 and GPT-5.2-Codex training stacks. The result was a single model that could do both deep software engineering and general knowledge work without a context switch. OpenAI described the move as Codex going "from an agent that can write and review code to an agent that can do nearly anything developers and professionals can do on a computer." [1]

A notable detail in the launch announcement: early prototypes of the model were used during its own development. The Codex engineering team deployed prototypes to debug the model's own training, manage its deployment, and diagnose test results and evaluations. [1] This made GPT-5.3-Codex the first model OpenAI publicly described as materially contributing to its own production. In the company's own framing, it was "the first model that was instrumental in creating itself." [1]

The model accepts text and image input and produces text output. [25] It runs on the Responses, Chat Completions, Assistants, Batch, and Realtime API endpoints. Fine-tuning was not enabled at launch. The knowledge cutoff is August 31, 2025, the same as GPT-5.2 and most of the rest of the GPT-5 series. [25][26]

### GPT-5.3-Codex-Spark

GPT-5.3-Codex-Spark launched on February 12, 2026, as a research preview available to ChatGPT Pro subscribers. [3] It is a distilled, smaller variant of GPT-5.3-Codex optimized for near-instant inference rather than deep reasoning. OpenAI developed it in partnership with [Cerebras](/wiki/cerebras) as the first milestone in a collaboration announced in January 2026. [3][22] Spark runs on the Cerebras Wafer-Scale Engine 3 (WSE-3), a purpose-built AI accelerator, and delivers more than 1,000 tokens per second on real-world coding tasks. [3][22][24] In an OpenAI demo, the Cerebras-backed Spark completed a "build a snake game" task in about 9 seconds, compared with nearly 43 seconds on the non-Spark model. [3]

The model is intended for inline code completions, boilerplate generation, quick refactors, and short-scope tasks where sub-second latency matters more than multi-step reasoning depth. At launch it supported text input only; image support was not included in the initial preview. [3][23] OpenAI described its release as an early access phase while Cerebras ramps up datacenter capacity to handle production demand. Access was rolled out progressively through the Codex app, the Codex CLI, and the Codex VS Code extension. API access was distributed to a small set of design partners rather than opened generally. [3]

Spark is also notable as the first OpenAI model not running on NVIDIA hardware, marking the company's first production deployment on silicon outside its core NVIDIA stack. [24] The Cerebras WSE-3 chip is fundamentally different in architecture: it is a single wafer-scale processor with very large on-chip SRAM, which allows the entire active model to live in fast memory rather than being streamed from HBM. [22][24] That architecture is what makes the throughput possible, but it also constrains the size of the model that can be deployed, which is why Spark is a distilled variant rather than the full Codex model.

### GPT-5.3 Instant

GPT-5.3 Instant is the conversational successor to GPT-5.2 Instant and the default model for all ChatGPT users from March 3, 2026. [4] Its most-discussed change was behavioral rather than a raw capability gain. OpenAI addressed a sustained wave of user complaints that GPT-5.2 Instant used moralizing, preachy language and unsolicited emotional coaching. Phrases like "Stop. Take a breath." became widely mocked on social media in the weeks following the GPT-5.2 Instant rollout. [4][13][16]

OpenAI summarized the update bluntly. In its launch communications the company wrote: "We heard your feedback loud and clear, and 5.3 Instant reduces the cringe." [16] The model was tuned to acknowledge difficulties without patronizing reassurance, answer questions that GPT-5.2 had refused unnecessarily, and cut defensive preambles and safety disclaimers from situations where they served no purpose. [4][14] OpenAI described the focus as improving "tone, relevance, and conversational flow," characteristics that may not show up in benchmark numbers but heavily affect whether users keep coming back. [4]

Beyond tone, GPT-5.3 Instant also expanded the context window from 200K tokens (GPT-5.2 Instant) to 400K tokens. [20] That enables single-call processing of roughly 300,000 words, which covers most book-length documents and large meeting transcripts in a single request.

## What are the GPT-5.3 technical specifications?

### GPT-5.3-Codex specifications

| Property | Value |
|---|---|
| Release date | February 5, 2026 [1] |
| Context window | 400K tokens [25] |
| Input modalities | Text, images [25] |
| Output modalities | Text [25] |
| API identifier | `gpt-5.3-codex` [5] |
| Knowledge cutoff | August 31, 2025 [25] |
| Reasoning effort levels | low, medium, high, xhigh [25] |
| Input pricing | $1.75 per 1M tokens [25][26] |
| Cached input pricing | $0.175 per 1M tokens [11] |
| Output pricing | $14.00 per 1M tokens [25][26] |
| Endpoints | Responses, Chat Completions, Assistants, Batch, Realtime |
| Fine-tuning | Not supported at launch |

The 400K-token context window was a notable expansion over the 200K of earlier Codex variants. [25] Combined with token-efficient tool use, that allows the model to absorb very large repositories without hitting context limits during multi-step refactors.

### GPT-5.3-Codex-Spark specifications

| Property | Value |
|---|---|
| Release date | February 12, 2026 (research preview) [3] |
| Input modalities | Text only (at launch) [3] |
| Hardware | Cerebras Wafer-Scale Engine 3 (WSE-3) [22] |
| Throughput | 1,000+ tokens per second [3][24] |
| Availability | ChatGPT Pro (research preview), select API design partners [3] |
| Surfaces | Codex app, Codex CLI, Codex VS Code extension [3] |
| Pricing | Not publicly announced at preview launch |

### GPT-5.3 Instant specifications

| Property | Value |
|---|---|
| Release date | March 3, 2026 [4] |
| Context window | 400K tokens [20] |
| Input modalities | Text, images |
| API identifier | `gpt-5.3-instant` [20] |
| Input pricing | $1.10 per 1M tokens [20] |
| Cached input pricing | $0.55 per 1M tokens [20] |
| Output pricing | $4.40 per 1M tokens [20] |

The API identifier follows OpenAI's convention of also pinning the chat-tier model behind a moving alias (`chat-latest`), so workloads automatically pick up minor revisions without explicit version updates. [27]

## How does GPT-5.3-Codex perform on benchmarks?

GPT-5.3-Codex posted the strongest gains on terminal and computer-use tasks, where earlier Codex models had lagged behind their general-reasoning counterparts. OpenAI cited four headline benchmarks at launch: SWE-Bench Pro, Terminal-Bench, OSWorld, and GDPval. [1]

| Benchmark | GPT-5.3-Codex | GPT-5.2-Codex |
|---|---|---|
| [SWE-Bench](/wiki/swe-bench) Pro Public | 56.8% | 56.4% [10][11] |
| SWE-Lancer IC Diamond | 81.4% | 76.0% [11] |
| Terminal-Bench 2.0 | 77.3% | 64.0% [11][26] |
| OSWorld-Verified | 64.7% | 38.2% [11][26] |
| Cybersecurity CTF | 77.6% | 67.4% [11] |

OpenAI noted that GPT-5.3-Codex achieves its SWE-Bench Pro scores using fewer output tokens than any prior model, which means lower effective cost per task on agentic workflows. [1] Several developer reviewers reported the model used "less than half" the tokens of GPT-5.2-Codex on the same tasks. [19] The OSWorld-Verified score of 64.7% nearly doubled the GPT-5.2-Codex result of 38.2%, closing a gap that prior Codex models had not come close to closing. [11]

The Terminal-Bench 2.0 improvement of 13.3 percentage points was among the largest single-generation gains on that benchmark at the time. [11] The score reflects the model's enhanced ability to handle complex CLI operations, chained tool calls, and stateful shell sessions. On structured knowledge work measured by the GDPval benchmark, one analysis reported a 70.9% win-or-tie rate against baseline systems. [21] On the Artificial Analysis Intelligence Index, which aggregates harder evaluations including [GPQA](/wiki/gpqa) Diamond, Humanity's Last Exam, [tau-bench](/wiki/tau-bench), and Terminal-Bench Hard, GPT-5.3-Codex (xhigh reasoning) scored 44 and ranked among the top models measured, with a reported 400K context window and an August 31, 2025 knowledge cutoff. [25] (Artificial Analysis rankings shift as new models are added.)

### GPT-5.3-Codex-Spark benchmark results

Independent analyses estimated GPT-5.3-Codex-Spark's Terminal-Bench 2.0 score at approximately 58.4%, compared to GPT-5.3-Codex's 77.3%. [23] The tradeoff is intentional. The model was pruned and optimized for throughput, with reasoning depth deliberately reduced to support sub-second response times. [3] For interactive editor use cases, where the model fires off short suggestions while the user is mid-keystroke, that kind of accuracy is competitive with most editor-integrated completion tools.

### GPT-5.3 Instant benchmark results

| Benchmark | GPT-5.3 Instant | GPT-5.2 Instant |
|---|---|---|
| MMLU-Pro | 84.1% | 82.6% [20] |
| MATH-500 | 92.3% | 89.1% [20] |
| [HumanEval](/wiki/humaneval) | 95.1% | 93.2% [20] |
| SWE-bench Verified | 64.7% | 61.3% [20] |
| SimpleQA hallucination rate | 6.1% | 8.4% [17][20] |

For higher-stakes evaluations in domains such as medicine, law, and finance, OpenAI reported that GPT-5.3 Instant reduced hallucination rates by 26.8% when using the web and 19.7% when relying solely on internal knowledge, compared to GPT-5.2 Instant. On its user-feedback evaluation, hallucinations decreased by 22.5% with web access and 9.6% without it. [4][17] OpenAI's published category breakdowns reported scientific fact retrieval improving by roughly 31.2% and medical fact retrieval by roughly 29.7%. [20]

OpenAI attributed the hallucination improvements largely to better calibration: earlier training had rewarded confident helpfulness over calibrated uncertainty, and the new model was tuned to express uncertainty proportionally to its actual reliability on a query rather than reaching for a confident-sounding fabrication. [17] The result was a model that was more willing to say it did not know something. The benchmark gains over GPT-5.2 Instant are best read as incremental rather than generational. Multiple reviewers characterized the academic changes as roughly "+1 to +3 points," with the headline number being the hallucination reduction rather than any single-benchmark leap. [20]

## How much does GPT-5.3 cost?

### API pricing

GPT-5.3-Codex sits at the same price point as its Codex predecessor: $1.75 per million input tokens and $14.00 per million output tokens. [25][26] Cached input tokens are priced at $0.175 per million, a 90% discount that rewards repeat-context workloads such as agentic loops working over the same repository state. [11] This positions GPT-5.3-Codex within the high-capability reasoning tier rather than as a low-cost inference option.

GPT-5.3 Instant is priced at $1.10 per million input tokens and $4.40 per million output tokens, with cached input at $0.55 per million. [20] That is a meaningful reduction from the GPT-5.2 series' higher-cost models, placing it as a competitive option in the fast-inference tier against models like Anthropic's Claude Haiku and Google's Gemini Flash. The pricing structure also widens the cost gap between Instant and Codex, which encourages routing decisions where simple chat traffic stays on Instant and only agentic workloads escalate to Codex.

GPT-5.3-Codex-Spark pricing was not publicly announced at the time of its research preview launch. [3] Access during the preview phase was included for ChatGPT Pro subscribers and select API design partners. The Cerebras infrastructure cost structure differs significantly from NVIDIA-based deployment, which made standard per-token comparisons less meaningful during the early access period.

### ChatGPT subscription access

| Plan | GPT-5.3 Instant | GPT-5.3-Codex | GPT-5.3-Codex-Spark |
|---|---|---|---|
| Free | Default model | No | No |
| Plus | Default model | Yes (via Codex surfaces) | No |
| Pro | Default model | Yes | Yes (research preview) |
| Enterprise / Edu | Available | Available | Available on request |

For most of February 2026, GPT-5.3-Codex was a paid-tier feature available across the Codex app, CLI, IDE extension, and web, with API access planned once it was safely enabled. [1] ChatGPT Pro subscribers received priority access to all three GPT-5.3 variants through the Codex app, CLI, and VS Code extension. [3]

## Why was GPT-5.3-Codex rated "High" for cybersecurity?

GPT-5.3-Codex was the first model OpenAI classified as High capability for cybersecurity-related tasks under its Preparedness Framework. [1][2] It was also the first OpenAI model trained specifically to identify software vulnerabilities, which makes the cybersecurity rating less surprising in retrospect than it sounded at announcement. [2]

Under the Preparedness Framework, the High cybersecurity threshold is defined as a model that removes existing bottlenecks to scaling cyber operations, including by automating end-to-end cyber operations against reasonably hardened targets, or by automating the discovery and exploitation of operationally relevant vulnerabilities. [2] The threshold is one of the higher bars in the framework, and crossing it triggered safety controls that had not been used for any previous OpenAI release. [1]

OpenAI stated in the system card that it does not have definitive evidence that GPT-5.3-Codex reaches the threshold, but adopted a precautionary approach because it could not rule out the possibility. [2] [Sam Altman](/wiki/sam_altman) confirmed the classification publicly, describing it as "our first model that hits 'high' for cybersecurity on our preparedness framework." [12] Altman framed the announcement carefully, emphasizing that the rating reflected what the model could potentially do rather than what it had been observed doing in deployment. [12]

The Cybersecurity CTF benchmark result of 77.6% is the primary empirical basis for the classification. [11] The model demonstrates breadth of end-to-end successes consistent with the High definition, including automated vulnerability discovery and exploitation in controlled evaluations. OpenAI stated in the system card that the same capabilities that make the model effective at writing, testing, and reasoning about code also raise serious cybersecurity concerns, and that the dual-use character of those capabilities is unusually pronounced for a coding model. [2]

In response, OpenAI said it was "deploying our most comprehensive cybersecurity safety stack to date," with mitigations including safety training, automated monitoring, trusted access for advanced capabilities, and enforcement pipelines including threat intelligence. [12] In practice this meant:

- The model is trained to refuse clearly malicious requests, including direct asks for offensive cyber tools. [12]
- Automated classifier-based monitors detect signals of suspicious cyber activity and route high-risk traffic to a less cyber-capable fallback model. [2]
- Full API access was deliberately delayed after the ChatGPT rollout to allow additional safety review. [1][12]
- A Trusted Access Program for vetted security professionals gates advanced capabilities for verified researchers and defenders. [6][12]
- OpenAI offered $10 million in API credits for cybersecurity defense applications, encouraging the model to flow toward defensive use cases. [12]
- Threat intelligence enforcement pipelines monitor for misuse patterns and feed back into model training and refusal policies. [12]

Fortune characterized the launch as one where OpenAI had "run headlong into the risks of releasing" a model that crossed a new cybersecurity risk threshold. [12] Some security researchers questioned whether the Trusted Access Program could prevent misuse at scale, particularly given that the same technical capabilities that enable defenders also enable adversaries. Others noted that the disclosure-and-mitigation pattern was a real-world test of the framework's operational design. [12]

## How does GPT-5.3 compare to GPT-5.2 and competitors?

### GPT-5.3-Codex vs. GPT-5.2-Codex

GPT-5.3-Codex is 25% faster than GPT-5.2-Codex on equivalent tasks, achieved through improvements to both the inference stack and the model architecture. [1][9] Its SWE-Bench Pro score is only marginally higher (+0.4 points), but its Terminal-Bench 2.0 score is 13.3 points higher and its OSWorld-Verified score is 26.5 points higher. [11] These improvements reflect a fundamental broadening of the model's agentic scope. Where GPT-5.2-Codex was strongest on repository-level code changes, GPT-5.3-Codex extends that capability to terminal-based operations and visual desktop environments. The token efficiency gain matters too: developers reported that the same task often used less than half the output tokens, which directly cuts cost on long agentic loops. [19]

### GPT-5.3-Codex vs. Claude Opus 4.6

The head-to-head comparison with Anthropic's [Claude Opus 4.7](/wiki/claude_opus_4_7) family (specifically Opus 4.6, which was current at the GPT-5.3-Codex launch) was a focal point of February 2026 developer coverage. [12] In hands-on reviews, GPT-5.3-Codex finished tasks faster than Claude Opus 4.6 and was described as harder to beat on well-specified tasks with clear validation criteria. Claude Opus 4.6 showed stronger first-attempt reliability and better performance on tasks with ambiguous or underspecified instructions. [19] Reviewers characterized the choice as speed and precision on known-good instructions (GPT-5.3-Codex) versus reliability and reasoning with ambiguity (Opus 4.6).

One developer described GPT-5.3-Codex as the first model that made it plausible to "specify the outcome, set up validation with clear pass/fail tests, and press go." [19] That was a meaningful shift in agentic coding workflows, because the prior pattern had been to specify intermediate steps closely and supervise the model heavily. The two models settled into complementary roles in many teams, with Codex used for tightly-scoped agent runs and Opus used for exploratory work where the spec itself was being refined.

### GPT-5.3-Codex vs. Gemini 3

Comparisons with Google's [Gemini 3](/wiki/gemini_3) Pro family put GPT-5.3-Codex ahead on Terminal-Bench 2.0 and OSWorld-Verified by clear margins, while Gemini 3 retained an advantage on multimodal reasoning tasks involving long video and large image collections. On standard SWE-Bench Pro, the two models traded the lead depending on subtask category. Most developer evaluations published in February and March 2026 characterized the gap as task-dependent rather than uniform. [21]

### GPT-5.3 Instant vs. GPT-5.2 Instant

GPT-5.3 Instant made the most visible change in tone and behavioral calibration rather than raw capability. [4] The context window doubled from 200K to 400K tokens. [20] Benchmark improvements were incremental: MATH-500 rose from 89.1% to 92.3%, HumanEval from 93.2% to 95.1%, and SWE-bench Verified from 61.3% to 64.7%. [20] The hallucination rate reduction of 26.8% on web-enabled queries was the largest single quantified improvement OpenAI highlighted at launch. [17] OpenAI also cited internal blind A/B testing showing roughly a 34% improvement in user preference for the GPT-5.3 Instant tone over GPT-5.2 Instant, although the company did not publish methodology details. [20]

## What can GPT-5.3 be used for?

### GPT-5.3-Codex use cases

GPT-5.3-Codex is suited for multi-step, long-horizon software engineering tasks. That includes large refactoring projects spanning multiple files, complex debugging sessions requiring iterative diagnosis, architecture design, security review, and production monitoring. Its expanded OSWorld-Verified performance makes it usable for computer-use workflows involving visual desktop environments, not just CLI and API-based operations. [1][11]

OpenAI positioned the model as extending beyond pure coding into general knowledge work. A representative example in the launch announcement: the model could write a SQL query, fetch the data, then generate a PDF report or slide deck from the results, handling the full chain through tool calls. [1] That same chain previously required two separate model invocations or a hand-coded orchestrator.

The site reliability use case from OpenAI's own infrastructure deployment is also instructive. The model was used to monitor training runs, classify infrastructure errors, identify root causes for cache hit rate regressions, and trigger scaling actions on GPU clusters. [1] Those are open-ended, observation-driven tasks that combine reasoning over logs, knowledge of the system, and decisions about when to act. The fact that OpenAI shipped a production-grade workflow built on the model is a strong validation, although it also raises questions about whether internal tooling generalizes to outside teams who do not have OpenAI's specific instrumentation.

### GPT-5.3-Codex-Spark use cases

GPT-5.3-Codex-Spark is designed for developer workflow integration where latency directly affects productivity. That covers inline code completions in editors, real-time suggestions during live coding, unit test scaffolding, boilerplate generation, and syntax repair. [3] Its 1,000+ token-per-second throughput means suggestions appear before a developer has finished typing the context, which reduces the interruption cost of switching cognitive modes between writing and reviewing. [3][24]

It is not suited for tasks requiring extended reasoning chains, multi-file analysis, or complex architectural decisions. For those, the standard GPT-5.3-Codex remains the recommended option. The intended split is that Spark covers the moment-by-moment rhythm of typing while Codex handles the longer agent runs, similar to the way some IDE setups already mix a small local completion model with a larger remote reasoning model.

### GPT-5.3 Instant use cases

GPT-5.3 Instant covers the full range of everyday ChatGPT workflows: writing assistance, research and summarization, translation, general question answering, image understanding, and web search integration. [4] Its hallucination reductions are most pronounced in scientific and medical domains, which makes it more reliable for healthcare and research queries than its predecessor without crossing into the territory of specialized verticalized models. [17][20]

The 400K token context window enables processing of book-length documents, large codebases, or extended conversation histories without truncation. [20] That has practical effects for legal, medical, and academic users who routinely paste in long source documents and ask follow-up questions across the entire content. Image understanding remains supported, which keeps GPT-5.3 Instant viable for visual question answering and OCR-adjacent workflows.

## How was GPT-5.3 received?

### GPT-5.3-Codex reception

Developer reception for GPT-5.3-Codex was largely positive. Reviewers described it as a straight upgrade for existing users and noted the combination of speed gains, broader agentic scope, and improved interactive behavior during code reviews. [19] The model's ability to explain its changes, suggest alternatives, and adapt to feedback mid-session was called out as a meaningful improvement over GPT-5.2-Codex, which often completed tasks without commentary. Reviewers also highlighted the token efficiency, with one early-access review noting the model used "less than half" the tokens of GPT-5.2-Codex on the same tasks. [19]

Some developers noted persistent limitations. The model still interprets instructions literally rather than inferring intent when instructions are underspecified. Reviewers noted cases where the model ran extended tool call sequences without identifying the root issue. [18] Unexplained drops in mid-session quality were also reported, attributed by some developers to possible routing to lighter model variants under load.

The cybersecurity classification attracted significant attention from security researchers. Coverage in Fortune and other outlets described the model's risk posture as a step into new territory for a publicly accessible coding model. [12] Some security researchers questioned whether the Trusted Access Program was sufficient to prevent misuse at scale. Others welcomed the explicit safety framing, on the view that having a public Preparedness Framework rating was a meaningful improvement over earlier capability releases that left risk evaluation implicit. [12]

### GPT-5.3-Codex-Spark reception

GPT-5.3-Codex-Spark generated enthusiasm among developers who had found prior AI code completions too slow to integrate into their editing flow. The Cerebras partnership was covered as a notable infrastructure milestone, given that the WSE-3 wafer-scale chip architecture is fundamentally different from the NVIDIA GPU clusters used for most large model inference. [22][24] AI Business and other industry outlets covered the partnership as a signal that OpenAI was actively diversifying its inference hardware base, which had become a strategic concern as GPU supply tightened in 2025. [24]

At launch, availability was limited to ChatGPT Pro subscribers, with no firm date announced for broader rollout. [3] Some developers questioned whether the speed advantage was worth the reasoning tradeoff, particularly for teams already using IDE-integrated completion tools from other providers. Others praised the focused product positioning, noting that Spark did one thing well rather than trying to be a generalist. [23]

### GPT-5.3 Instant reception

GPT-5.3 Instant's release attracted more media attention for its behavioral changes than for its benchmark numbers. Headlines in TechCrunch ("ChatGPT's new GPT-5.3 Instant model will stop telling you to calm down"), TechRadar ("We heard your feedback loud and clear, OpenAI introduces new ChatGPT 5.3 Instant to reduce the cringe"), and Decrypt ("More accurate, less cringe") captured the user sentiment. [13][15][16] VentureBeat framed the launch as OpenAI shifting focus from speed to accuracy, arguing that the hallucination reductions were the more durable change behind the headline tone work. [17]

Reaction on social media was mixed. Many users praised the more direct, less patronizing style. Others took the occasion to raise broader criticisms of OpenAI's policy direction. Some users requested the return of [GPT-4o](/wiki/gpt_4o) or earlier ChatGPT defaults, citing preference for the tone of older models over any of the GPT-5 family. [13][14]

The "reduces the cringe" line itself drew commentary. Some observers read it as a candid acknowledgment that the previous tone had been a real product problem. Others saw it as marketing language that papered over the more interesting question of why a previous model release had been tuned in that direction in the first place. [16] Either way, it set a precedent for OpenAI public communications, and similar phrasing reappeared in later 2026 releases.

## What are the limitations of GPT-5.3?

### GPT-5.3-Codex limitations

Despite its improvements, GPT-5.3-Codex inherits known weaknesses from the Codex line. It performs best with precise, detailed instructions and degrades when tasks are underspecified. Unlike Claude Opus 4.6, it does not reliably infer developer intent when the specification is ambiguous. [19] Reviewers noted a tendency to get stuck in extended diagnostic loops on problems that could have been resolved earlier with broader context, although the larger 400K context window mitigates this in many real-world cases. [18]

The model's knowledge cutoff of August 31, 2025 limits its awareness of libraries, frameworks, and APIs released or updated after that date. [25] The cybersecurity risk profile requires OpenAI to maintain active monitoring infrastructure, which adds latency overhead for certain query categories routed through the automated classifier stack. Those routing decisions are not always visible to developers, which can produce occasional unexplained latency or quality variation on sensitive prompts. [2]

### GPT-5.3-Codex-Spark limitations

GPT-5.3-Codex-Spark trades reasoning depth for speed. Its estimated Terminal-Bench 2.0 score is roughly 19 percentage points below the standard Codex, and it is not suited for multi-step reasoning or complex debugging. [23] Text-only input at launch excluded image and screenshot workflows, which are common in editor-integrated debugging. [3]

Availability was restricted to ChatGPT Pro at launch, with no public API access announced beyond a small set of design partners. [3] The Cerebras hardware dependency also creates a different scaling profile from NVIDIA-based inference, which means availability is bounded by Cerebras datacenter buildout rather than the more elastic GPU pools that power the rest of the OpenAI lineup. [24]

### GPT-5.3 Instant limitations

GPT-5.3 Instant is a general-purpose conversational model and does not match the deep reasoning capability of GPT-5.2 Thinking or the specialized coding performance of GPT-5.3-Codex. Its benchmark gains over GPT-5.2 Instant are incremental rather than transformative. [20] Multiple reviewers noted that on long single-pass reasoning problems, the Thinking-tier model is still the right choice.

The anti-cringe tuning reduced some forms of excessive caution but did not eliminate all refusals. The model still declines categories of requests that some users consider legitimate, and the balance between helpfulness and safety remains a subject of ongoing user feedback. [16] There is also a more subtle concern about behavioral oscillation. Some users who had adjusted their workflows to GPT-5.2 Instant's style reported that the new tone felt unfamiliar in the first week, which is the kind of mid-cycle adjustment cost that any conversational tone change introduces.

## What came after GPT-5.3?

GPT-5.3 Instant remained the ChatGPT default for two months. On May 5, 2026, OpenAI released GPT-5.5 Instant, which replaced GPT-5.3 Instant as the default model for all users. [27] GPT-5.5 Instant reported 52.5% fewer hallucinated claims than GPT-5.3 Instant on high-stakes prompts in medicine, law, and finance, and shipped tighter, less emoji-heavy, more personalized responses. [27] GPT-5.3 Instant remained available to paid ChatGPT subscribers for a three-month transition period after the GPT-5.5 default switch. [27]

GPT-5.3-Codex remained the production Codex model into the second quarter of 2026 and continued to receive minor updates, including expanded reasoning depth options and improved tool-use prompts in the Responses API. GPT-5.3-Codex-Spark gradually expanded availability beyond the initial Pro-only research preview as Cerebras infrastructure scaled, although the model remained more of a specialty offering than a general-purpose default. [3]

The broader release cadence of February through May 2026 (Codex, Codex-Spark, Instant, then the [GPT-5.4](/wiki/gpt-5.4) family) cemented OpenAI's pattern of fast minor-version updates, with the version-and-suffix naming scheme starting to strain. By mid-2026 the company was managing a model lineup that included several active versions of Instant, Thinking, Codex, and Pro, plus partner-specific variants like Codex-Spark, all under the GPT-5 umbrella. That complexity drew its own commentary, with developer outlets running explainer guides on which model to pick for which task. [21]

## ELI5: GPT-5.3 in simple terms

Think of GPT-5.3 as three new versions of OpenAI's AI that came out around the same time in early 2026. One (Codex) is a super-smart coding helper that can do almost anything on a computer, like writing software and fixing problems, and it was so capable at finding security holes that OpenAI added extra safety locks. [1][2] Another (Codex-Spark) is a stripped-down, lightning-fast version that finishes simple coding tasks almost instantly, running on a giant special chip from a company called Cerebras instead of the usual NVIDIA chips. [3][24] The third (Instant) is the everyday chatbot in ChatGPT, and its big update was mostly about being less preachy and annoying. OpenAI literally said the goal was to "reduce the cringe" and make it make fewer mistakes. [16][17]

## See also

- [GPT-5](/wiki/gpt-5)
- [GPT-5.1](/wiki/gpt-5.1)
- [GPT-5.2](/wiki/gpt-5.2)
- [GPT-5.4](/wiki/gpt-5.4)
- [GPT-5.5](/wiki/gpt-5.5)
- [GPT-4](/wiki/gpt-4)
- [GPT-4.1](/wiki/gpt-4.1)
- [o1](/wiki/o1)
- [o3](/wiki/o3)
- [OpenAI Codex](/wiki/codex)
- [ChatGPT](/wiki/chatgpt)
- [OpenAI](/wiki/openai)
- [OpenAI API](/wiki/openai_api)
- [Sam Altman](/wiki/sam_altman)
- [Cerebras](/wiki/cerebras)

## References

1. OpenAI. "Introducing GPT-5.3-Codex." openai.com, February 5, 2026. https://openai.com/index/introducing-gpt-5-3-codex/
2. OpenAI. "GPT-5.3-Codex System Card." cdn.openai.com, February 5, 2026. https://cdn.openai.com/pdf/23eca107-a9b1-4d2c-b156-7deb4fbc697c/GPT-5-3-Codex-System-Card-02.pdf
3. OpenAI. "Introducing GPT-5.3-Codex-Spark." openai.com, February 12, 2026. https://openai.com/index/introducing-gpt-5-3-codex-spark/
4. OpenAI. "GPT-5.3 Instant: Smoother, more useful everyday conversations." openai.com, March 3, 2026. https://openai.com/index/gpt-5-3-instant/
5. OpenAI. "GPT-5.3-Codex Model." developers.openai.com. https://developers.openai.com/api/docs/models/gpt-5.3-codex
6. OpenAI. "Introducing Trusted Access for Cyber." openai.com. https://openai.com/index/trusted-access-for-cyber/
7. OpenAI. "Model Release Notes." help.openai.com. https://help.openai.com/en/articles/9624314-model-release-notes
8. Wikipedia. "GPT-5.3-Codex." en.wikipedia.org. https://en.wikipedia.org/wiki/GPT-5.3-Codex
9. Neowin. "OpenAI debuts GPT-5.3-Codex: 25% faster and setting new coding benchmark records." neowin.net, February 5, 2026. https://www.neowin.net/news/openai-debuts-gpt-53-codex-25-faster-and-setting-new-coding-benchmark-records/
10. DataCamp. "GPT-5.3 Codex: From Coding Assistant to General Work Agent." datacamp.com. https://www.datacamp.com/blog/gpt-5-3-codex
11. Digital Applied. "GPT-5.3 Codex: Features, Benchmarks, and Migration Guide." digitalapplied.com. https://www.digitalapplied.com/blog/gpt-5-3-codex-release-features-benchmarks-guide
12. Fortune. "OpenAI's new model leaps ahead in coding capabilities but raises unprecedented cybersecurity risks." fortune.com, February 5, 2026. https://fortune.com/2026/02/05/openai-gpt-5-3-codex-warns-unprecedented-cybersecurity-risks/
13. TechCrunch. "ChatGPT's new GPT-5.3 Instant model will stop telling you to calm down." techcrunch.com, March 3, 2026. https://techcrunch.com/2026/03/03/chatgpts-new-gpt-5-3-instant-model-will-stop-telling-you-to-calm-down/
14. 9to5Mac. "OpenAI releases GPT-5.3 Instant update to make ChatGPT less 'cringe'." 9to5mac.com, March 3, 2026. https://9to5mac.com/2026/03/03/openai-releases-gpt-5-3-instant-update-to-make-chatgpt-less-cringe/
15. Decrypt. "'More Accurate, Less Cringe': OpenAI Rolls Out GPT-5.3 Instant in ChatGPT." decrypt.co, March 3, 2026. https://decrypt.co/359837/more-accurate-less-cringe-openai-gpt-5-3-instant-chatgpt
16. TechRadar. "'We heard your feedback loud and clear', OpenAI introduces new ChatGPT 5.3 Instant to 'reduce the cringe'." techradar.com, March 3, 2026. https://www.techradar.com/ai-platforms-assistants/chatgpt/we-heard-your-feedback-loud-and-clear-openai-introduces-new-chatgpt-5-3-instant-to-reduce-the-cringe-for-all-users
17. VentureBeat. "GPT-5.3 Instant cuts hallucinations by 26.8% as OpenAI shifts focus from speed to accuracy." venturebeat.com. https://venturebeat.com/orchestration/gpt-5-3-instant-cuts-hallucinations-by-26-8-as-openai-shifts-focus-from
18. Golchian, Pooya. "GPT-5.3-Codex Performance Analysis: SWE-Bench Pro, Terminal-Bench 2.0, OSWorld Results." pooya.blog, 2026. https://pooya.blog/blog/gpt-5-3-codex-autonomous-coding-agent-2026/
19. Every. "Vibe Check: GPT-5.3 Codex, The 10x Engineer, Now More Fun at Parties." every.to. https://every.to/vibe-check/gpt-5-3-codex
20. Digital Applied. "GPT-5.3 Instant: Benchmarks, Pricing, Migration." digitalapplied.com. https://www.digitalapplied.com/blog/gpt-5-3-instant-anti-cringe-update-benchmarks-pricing
21. NxCode. "OpenAI GPT-5 Model Guide: GPT-5.2 vs 5.3 vs 5.4, Which One Should You Use? (2026)." nxcode.io. https://www.nxcode.io/resources/news/openai-gpt-5-model-guide-which-to-use-2026
22. Cerebras. "Introducing OpenAI GPT-5.3-Codex-Spark Powered by Cerebras." cerebras.ai. https://www.cerebras.ai/blog/openai-codexspark
23. Medium (CometAPI). "GPT-5.3 Codex Spark vs GPT-5.3 Codex: Comprehensive analysis." medium.com. https://medium.com/@mkteam/gpt-5-3-codex-spark-vs-gpt-5-3-codex-comprehensive-analysis-9a3f9fb00041
24. AI Business. "OpenAI GPT-5.3-Codex-Spark Shows What's Possible With Cerebras." aibusiness.com. https://aibusiness.com/generative-ai/openai-gpt-5-3-codex-spark-shows-what-s-possible-with-cerebras
25. Artificial Analysis. "GPT-5.3 Codex (xhigh) Intelligence, Performance & Price Analysis." artificialanalysis.ai. https://artificialanalysis.ai/models/gpt-5-3-codex
26. llm-stats. "GPT-5.3 Codex Benchmarks, Pricing & Context Window." llm-stats.com. https://llm-stats.com/models/gpt-5.3-codex
27. TechCrunch. "OpenAI releases GPT-5.5 Instant, a new default model for ChatGPT." techcrunch.com, May 5, 2026. https://techcrunch.com/2026/05/05/openai-releases-gpt-5-5-instant-a-new-default-model-for-chatgpt/

