# Claude Sonnet 5.5

> Source: https://aiwiki.ai/wiki/claude_sonnet_5_5
> Updated: 2026-09-29
> Fact-checked: 2026-09-29
> Categories: AI Models, AI Safety, Anthropic, Large Language Models
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Claude Sonnet 5.5." aiwiki.ai, 29 Sept 2026. https://aiwiki.ai/wiki/claude_sonnet_5_5
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

**Claude Sonnet 5.5** is a [large language model](https://aiwiki.ai/wiki/large_language_model) developed by [Anthropic](https://aiwiki.ai/wiki/anthropic) and released on September 28, 2026. Anthropic calls it "the second model in the Claude 5.5 family," after [Claude Opus 5.5](https://aiwiki.ai/wiki/claude_opus_5_5) (September 22, 2026), and it succeeds [Claude Sonnet 5](https://aiwiki.ai/wiki/claude_sonnet_5) in the mid-priced Sonnet tier of the [Claude](https://aiwiki.ai/wiki/claude) lineup.[1][15] Anthropic says it "runs 30%+ faster, and costs up to 30% less for most work" than Sonnet 5, while keeping Sonnet 5's list price of $2 per million input tokens and $10 per million output tokens.[1] The company positions it as "a faster, lower-cost complement" to Opus 5.5: Opus 5.5 is meant for complex work that needs careful judgment, and Sonnet 5.5 for "well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets."[1]

Sonnet 5.5 is the first Sonnet model to launch with the cyber safeguards and model fallbacks Anthropic had used on its most capable models, and the first Sonnet with safety classifiers that block attempts to extract its reasoning. Anthropic tied the cyber safeguards to the model's cybersecurity capability, which it describes as "comparable to Opus 5's," and the reasoning-extraction classifiers to the model being "far more capable than its predecessor."[1] On Anthropic's own launch table, Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0, against 10.3% for Sonnet 5 and 66.4% for Opus 5.5.[1] The independent evaluator [Artificial Analysis](https://aiwiki.ai/wiki/artificial_analysis) placed it second on its Intelligence Index on launch day, two points behind Opus 5.5, but measured the highest output-token use per task it had recorded.[22]

## Overview

| Property | Value |
| --- | --- |
| Developer | Anthropic |
| Release date | September 28, 2026 |
| Family | Claude 5.5 (second model released, after Opus 5.5) |
| Predecessor | [Claude Sonnet 5](https://aiwiki.ai/wiki/claude_sonnet_5) |
| Claude API model ID | `claude-sonnet-5-5` |
| [Amazon Bedrock](https://aiwiki.ai/wiki/amazon_bedrock) model ID | `anthropic.claude-sonnet-5-5` |
| Google Cloud, Microsoft Foundry, Claude Platform on AWS model ID | `claude-sonnet-5-5` |
| [Context window](https://aiwiki.ai/wiki/context_window) | 1M tokens |
| Max output | 128K tokens (300K on the Message Batches API with the `output-300k-2026-03-24` beta header) |
| Input / output | Text and images in, text out |
| Reliable knowledge cutoff | June 2026 |
| Training data cutoff | June 2026 |
| Thinking | [Adaptive thinking](https://aiwiki.ai/wiki/adaptive_thinking), on by default; lowest setting is `between_tools` |
| Effort levels | `low`, `medium`, `high`, `xhigh`, `max` |
| Default effort | `high` on the Claude API; `medium` in Claude Code and the Claude apps |
| Price (input / output) | $2 / $10 per million tokens |
| Comparative latency | "Fast" |
| Retirement | Not sooner than September 28, 2027 |

Sources: Anthropic model documentation, launch post and developer blog.[1][3][13][36]

## Background and positioning

Anthropic's main Claude models come in three sizes: Haiku (smallest and cheapest), Sonnet (medium) and Opus (largest of the three), with the more capable and more expensive Fable and Mythos models above them.[33] Sonnet 5 was released on June 30, 2026, and Opus 5.5 opened the Claude 5.5 family six days before Sonnet 5.5.[15] At the Opus 5.5 launch Anthropic had said that "Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks."[35] With Sonnet 5.5, Anthropic again said that Claude Haiku 5.5, "built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks."[1] As of September 29, 2026, Haiku 5.5 had been announced but not released, and Anthropic's model comparison still listed [Claude Haiku 4.5](https://aiwiki.ai/wiki/claude_haiku_4_5) as its fastest current model.[36]

Anthropic's documentation compares the current models as follows:[3]

| Model | Context | Max output | Price (input / output per MTok) | Latency | Default effort | Knowledge cutoff |
| --- | --- | --- | --- | --- | --- | --- |
| [Claude Fable 5.1](https://aiwiki.ai/wiki/claude_fable_5_1) | 1M | 128K | $10 / $50 | Slower | high | Jun 2026 |
| [Claude Opus 5.5](https://aiwiki.ai/wiki/claude_opus_5_5) | 1M | 128K | $4 / $20 | Moderate | medium | Jun 2026 |
| Claude Sonnet 5.5 | 1M | 128K | $2 / $10 | Fast | high | Jun 2026 |
| Claude Haiku 4.5 | 200K | 64K | $1 / $5 | Fastest | n/a | Feb 2025 |

The launch post says Sonnet 5.5 "doesn't advance the frontier of our models' capabilities." It adds that on several evaluations Sonnet 5.5 at Max effort "even performs comparably to Opus 5.5," but that "Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment," both in Anthropic's testing and in that of external testers.[1] Anthropic's developer blog puts the division this way: Sonnet 5.5 "fits best when the task has a clear spec and a way to check the result."[13] An Anthropic research product manager, Theo Chu, told CNBC that "Sonnet is really for the cost-conscious customer where they might not need as much intelligence."[29]

Anthropic lists five areas of improvement over Sonnet 5: performance, collaboration, cost, speed, and alignment and safety. It says Sonnet 5.5 writes more clearly than the previous generation, as Opus 5.5 does, that early testers described it as "a better partner for collaboration than Sonnet 5," and that it is "our fastest Sonnet model to date." The company also says it is "the first Sonnet model to beat Pokémon Red working only from screenshots."[1]

## Capabilities and benchmarks

### Launch benchmark table

The launch post compares Sonnet 5.5 with Sonnet 5, Opus 5.5 and OpenAI's [GPT-6 Sol](https://aiwiki.ai/wiki/gpt_6_sol). Anthropic did not print a GPT-6 Sol figure for four rows.[1]

| Benchmark (area) | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
| --- | --- | --- | --- | --- |
| [Terminal-Bench](https://aiwiki.ai/wiki/terminal_bench) 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% (note 1) | not shown |
| FrontierCode 1.1, Main (agentic coding) | 46.2% at Max (note 2); 52.1% at Xhigh | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 (agentic coding) | 55.5% | 34.1% | 57.8% | not shown |
| [GDPval](https://aiwiki.ai/wiki/gdpval)-AA v2.1 (knowledge work, Elo; note 3) | 1844 | 1449 | 1846 | 1487 (note 4) |
| AA-Briefcase v1.1 (knowledge work, Elo; note 3) | 1811 | 1359 | 1822 | 1483 (note 4) |
| [Humanity's Last Exam](https://aiwiki.ai/wiki/humanity_s_last_exam) (multidisciplinary reasoning, with tools) | 64.5% | 54.9% | 67.7% | not shown |
| [OSWorld](https://aiwiki.ai/wiki/osworld) 2.1 (computer use, partial scoring) | 80.1% | 57.0% | 81.8% | not shown |
| Chartography (visual chart recognition, no tools) | 61.6% | 15.6% | 64.4% | 53.6% (note 4) |

Source: Anthropic launch post and its footnotes.[1]

The footnotes to the table read, in summary:[1]

1. The Opus 5.5 Terminal-Bench 4.0 figure is at Xhigh effort, "which represents the model's highest score."
2. Sonnet 5.5 scored lower at Max effort than at Xhigh on FrontierCode, which evaluates whether a code change could be merged without human edits and penalizes out-of-scope changes "even if they are high-quality or helpful." At Max, Sonnet 5.5 more often ran [Claude Code](https://aiwiki.ai/wiki/claude_code)'s code-review skill, which splits a review across many subagents; in two cases [Cognition](https://aiwiki.ai/wiki/cognition_ai) examined, this led to a timeout or to edits beyond the task's scope.
3. Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment on the Claude Platform that had a bug which "could degrade responses to requests that use structured outputs." Anthropic expects the effect "to be small and to understate its performance," and says the bug has since been fixed.
4. OpenAI had recently fixed a bug that degraded image understanding in GPT-6 Sol, and the Artificial Analysis and [Surge AI](https://aiwiki.ai/wiki/surge_ai) (Chartography) figures may not yet reflect the fix. Artificial Analysis does not expect major effects on its two benchmarks, and Anthropic's internal testing suggested the Chartography score was not affected.

On two of its cost charts (Terminal-Bench 4.0 and CursorBench 4.0), Anthropic plotted [GPT-5.6](https://aiwiki.ai/wiki/gpt_5_6) Sol instead of GPT-6 Sol, because GPT-6 Sol results for those benchmarks had not been reported publicly.[1] The system card adds that the public Terminal-Bench 4.0 leaderboard lists GPT-5.6 Sol at 37.3% and that OpenAI reported 57.9% for [GPT-6 Astra](https://aiwiki.ai/wiki/gpt_6_astra) at high effort.[2]

The system card gives more detail on the Terminal-Bench run. Sonnet 5.5 scored 70.6% with safeguards enabled; a fallback model answered 1.2% of requests, affecting 1.5% of trials. The run used Claude Code in `--bare` mode at max effort with no internet egress, averaged five trials per task over the benchmark's 66 tasks, and has a standard error of about ±2.5 points. Opus 5.5 scored 64.8% at max effort, which Anthropic calls within noise of its 66.4% at xhigh.[2]

### Additional system card results

The system card's capability summary uses adaptive thinking at max effort, averaged over five trials, unless noted. It adds these rows to the launch table:[2]

| Evaluation | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
| --- | --- | --- | --- | --- |
| [SWE-bench Pro](https://aiwiki.ai/wiki/swe_bench_pro) | 81.3 | 63.2 | 89.9 | n/a |
| [SWE-bench Multilingual](https://aiwiki.ai/wiki/swe_bench_multilingual) | 90.3 | 78.3 | 93.9 | n/a |
| SWE-bench Multimodal | 54.3 | 28.1 | 61.4 | n/a |
| Humanity's Last Exam (no tools) | 56.9 | 43.1 | 64.4 | n/a |
| [HealthBench](https://aiwiki.ai/wiki/healthbench) Professional (length-adjusted) | 69.2 | 57.8 | 65.6 | n/a |
| AutomationBench ([Zapier](https://aiwiki.ai/wiki/zapier)) | 44.7 | 10.7 | 42.5 | 32.0 |

For AutomationBench, Zapier ran Sonnet 5.5 at max effort with the API's default fallbacks enabled. The 42.5% for Opus 5.5 comes from re-running its refused tasks with fallbacks enabled; the Opus 5.5 system card had reported 40.0%.[2]

Other results in the system card, all at max effort unless stated:[2]

| Evaluation | Sonnet 5.5 | Comparison |
| --- | --- | --- |
| DeepSWE v1.1 (113 long-horizon software tasks) | 71.0% | n/a |
| FrontierCode v1.1 Extended | 64.4% at xhigh; 59.1% at max | n/a |
| FrontierSWE v2 (Proximal, 34 tasks) | 61.9% | GPT-6 Astra 65.5%, Opus 5.5 62.3%, Fable 5.1 56.3%, GPT-5.6 Sol 32.2% |
| Terminal-Bench-Science 0.1 | 59.9% | Opus 5.5 58.7%, Fable 5.1 52.6%, Opus 5 29.0% |
| CursorBench 4.0 by effort (measured by Cursor) | 39.2% (medium), 47.8% (high), 53.1% (xhigh), 55.5% (max) | n/a |
| ArXivMath (August 2026), no tools / with tools | 86.8% / 95.2% | Opus 5.5 91.2% / 96.9%, Fable 5.1 82.9% / 92.1% |
| ProgramBench (166 tasks) | 79.7% | Sonnet 5 77.3%, Opus 5.5 91.2% |
| Chartography with tools | 90.2% | Opus 5.5 89.0%, Fable 5.1 88.4% |
| BenchCAD Vision2Code (voxel IoU), no tools / with tools | 0.747 / 0.963 | Opus 5.5 0.730 / 0.962 |
| OSWorld 2.1 strict pass rate | 43.5% | Opus 5.5 48.7%, Sonnet 5 25.6% |
| OfficeQA / OfficeQA Pro | 76.9% / 65.6% | Opus 5.5 78.9% / 67.7%, Sonnet 5 75.1% / 62.1% |
| Toolathlon-Verified, Pass@1 | 77.8% | Opus 5.5 77.8%, Sonnet 5 74.7% |
| Legal Agent Benchmark, held-out set (high effort) | 11.7% all-pass; 92.1% mean criterion pass | n/a |
| PhysicianBench | 63.2% | Sonnet 5 37.4%, Opus 5.5 68.4% |
| Global MMLU (42 languages) | 92.1% | Opus 5.5 94.3%, Sonnet 5 89.2% |
| MILU (11 languages) | 91.6% | Opus 5.5 93.1%, Sonnet 5 89.3% |

Healthcare is one area where the system card reports Sonnet 5.5 ahead of Opus 5.5. On HealthBench Professional at max effort, the raw scores of Sonnet 5.5 and Opus 5.5 were both 77.1%. After the length adjustment that OpenAI, the benchmark's developer, now applies, Sonnet 5.5 scored 69.2% against 65.6% for Opus 5.5, with GPT-6 Astra highest at 70.3%.[2] On GDPval-AA, Sonnet 5.5 scored 1725 at xhigh effort while using about 67% fewer output tokens than at max; on AA-Briefcase it scored 1746 at xhigh with about 61% fewer output tokens.[2]

### Cost-per-task and effort claims

Anthropic's launch charts plot score against cost per task at each effort level. It says that "on several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task," and that it complements Opus 5.5 best at lower effort settings, while "at higher settings, it can perform comparably at a similar cost."[1] The specific claims are:[1]

| Benchmark | Anthropic's claim |
| --- | --- |
| Terminal-Bench 4.0 | At Medium effort, "far exceeds Sonnet 5's best score for less than a tenth of the cost per task" |
| FrontierCode v1.1 Main | At High effort, "matches GPT-6 Sol's best score for about a fifth of the cost per task"; 10 points above Sonnet 5 at the same setting "at about one fifteenth of the cost per task" |
| CursorBench 4.0 | At Low effort, exceeds Sonnet 5's best score "for less than a tenth of the cost per task"; best score within about two points of Opus 5.5 |
| AA-Briefcase v1.1 | At Medium effort, "bests Sonnet 5's best score for about one ninth of the cost per task" |

Default effort differs by product. Anthropic says: "In Claude Code and our apps, the default effort is set to Medium, while the Claude Platform defaults to High."[1] The API documentation says the effort levels are "recalibrated," so a given level does not produce the same amount of thinking as on Sonnet 5. It recommends starting at `high` unless a workload is agentic or latency-sensitive, at `medium` for well-specified agentic coding and multistep tool use, and at `medium` or `low` for chat and latency-sensitive work, using `xhigh` or `max` "only where your evals show a quality gain."[7] Anthropic's developer blog warns that at `xhigh` or `max` the model "will think longer and cost more," and that on some tasks "you may lose some of what makes Sonnet useful: its balance of quality, speed, and cost."[13] Cursor's documentation reports that on CursorBench, max effort "costs about 6x more per task than the default high effort."[20]

### Early-access reports

Anthropic published statements from early testers. The figures in them are the companies' own measurements.[1]

| Organization | Reported result |
| --- | --- |
| [Base44](https://aiwiki.ai/wiki/base44) | Across 118 real app builds, apps "scored level with Opus 5," in 3.6 iterations per build on average against 7.7 for Opus 5 |
| Balyasny Asset Management | On 2,441 finance tasks, ahead of Sonnet 5 at about 121k tokens per answer against 497k |
| Box | More accurate than the previous model, "2.4x faster," with 12% fewer total tokens |
| [Slack](https://aiwiki.ai/wiki/slack) | Better than Sonnet 5 on "almost all" offline Slackbot evals, with about 14% fewer output tokens |
| Zendesk | Tickets processed 20% faster on hundreds of real support cases |
| Unity | Completed 90% of tasks in its multi-step Unity Editor and coding benchmark |
| [Lovable](https://aiwiki.ai/wiki/lovable_ai) | "A third fewer tool calls and roughly half the shell runs to finish a task" |
| Atlassian | Rovo Agents run "up to 30% faster than they could with Sonnet 5" |

Epic Games chief operating officer Daniel Vogel said that "in Epic's early testing, Claude Sonnet 5.5 cleared the same quality bar you'd expect from a higher-tier model." David Loker, vice president of AI at [CodeRabbit](https://aiwiki.ai/wiki/coderabbit), said: "Sonnet 5's tendency to reach for web search too often and its high token use are both gone in this new model." Sualeh Asif, listed by Anthropic as director of ML at "SpaceXAI," said Sonnet 5.5 "delivers frontier-level performance on CursorBench 4.0 at 55.5%, second only to Opus 5.5." Kevin Ngo of Creator said: "When Claude Opus 5.5 sets the architecture and general framework for a game, I would feel confident in letting Sonnet 5.5 implement it."[1]

Anthropic also says that in head-to-head runs the model "batched tool calls together more than Sonnet 5, leading to fewer steps and lower costs." In an internal test, it gave the model a public company's quarterly earnings materials, call transcripts and a slide template, and asked for a 10-slide operating review; two experts judged the first draft "ready to send as is."[1]

## Pricing and speed

| Price per million tokens | Claude Sonnet 5.5 | Claude Opus 5.5 |
| --- | --- | --- |
| Input | $2 | $4 |
| Output | $10 | $20 |
| 5-minute cache write | $2.50 | $5 |
| 1-hour cache write | $4 | $8 |
| Cache read (hit) | $0.20 | $0.20 |
| Batch API input / output | $1 / $5 | $2 / $10 |

Source: Anthropic launch post and pricing documentation.[1][6]

Anthropic says Sonnet 5.5 is "priced the same as Sonnet 5," including prompt caching and batch rates, but "typically needs far fewer tokens to do the same work," so that "in our testing, it costs up to 30% less per task than its predecessor."[1][4] It uses the same tokenizer as Sonnet 5, so the same text produces the same token count.[4] US-only inference through the `inference_geo` parameter costs 1.1 times the standard rates.[6][13] Sonnet 5.5's list prices are half those of Opus 5.5 for input and output.[1][29] Anthropic also says Sonnet 5.5 "generates outputs 30%+ faster than Sonnet 5."[1] The developer blog says Sonnet 5.5's rate limits are separate from Sonnet 5's, with the same default tier values, and that it uses a high-resolution image tier of up to 2,576 pixels on the long edge, so a 2000×1500 image costs about 2.5 times as many tokens as on Sonnet 4.6, Sonnet 4.5 or Haiku 4.5.[13]

## API changes

The documentation lists five breaking changes for code moving from Sonnet 5:[3][4][5]

| Change | Detail |
| --- | --- |
| Thinking cannot be set to `disabled` | `thinking: {"type": "disabled"}` returns a 400 error. The new lowest setting, `between_tools`, turns off up-front thinking; it works only at `low`, `medium` and `high` effort and takes no other fields |
| No forced tool use | `tool_choice` of `any` or a named tool returns a 400 error; `auto` (with `strict: true` tools for schema-valid input) and `none` still work |
| Thinking blocks tied to model and conversation | Sonnet 5.5 reads thinking blocks from Sonnet 5, Opus 4.8, Haiku 4.5 and earlier models, but not from Opus 5, Opus 5.5 or any Fable or Mythos model. No other model reads Sonnet 5.5's blocks |
| Older computer-use tool rejected | On the Claude API and Google Cloud only the `computer_toolset_20260801` toolset is accepted; Amazon Bedrock still accepts `computer_20251124` |
| Fewer advisor pairings | With the advisor tool, a Sonnet 5.5 executor rejects Opus 4.8, Opus 4.7 and Sonnet 5 as advisors, and every accepted advisor returns its advice encrypted |

A sixth change alters the response without failing any request: notes longer than a sentence or two that the model writes between tool calls now arrive as "progress-update" `thinking` blocks, which are empty at the default display setting, so an interface that streams those notes goes quiet between tool calls unless it sets a `display` value or uses `between_tools`.[4] Anthropic's launch post highlighted the `between_tools` change: developers who ran Sonnet 5 with thinking off "need to switch to the new between_tools setting, which keeps up-front thinking off."[1] On Amazon Bedrock, structured outputs, including strict tool use, are not available for Sonnet 5.5.[5]

Sonnet 5.5 also supports per-message effort (beta), mid-conversation system messages and mid-conversation tool changes (beta), none of which are available on Sonnet 5. The minimum cacheable prompt fell to 512 tokens from 1,024 on Sonnet 5. Non-default sampling parameters (`temperature`, `top_p`, `top_k`) and assistant prefill return errors, as they do on Sonnet 5.[3][4][5]

## Safety and safeguards

### Responsible Scaling Policy determination

The system card describes Sonnet 5.5 as "broadly less capable than Opus 5.5 across domains" and says it "does not cross any new RSP thresholds." Because Anthropic had already determined that Opus 5.5 does not cross the CB-2 or Autonomy-2 thresholds of its [Responsible Scaling Policy](https://aiwiki.ai/wiki/responsible_scaling_policy), it applied the same conclusion to Sonnet 5.5. The company treats Sonnet 5.5 as meeting its CB-1 and Autonomy-1 thresholds and applies the corresponding mitigations. The card frames its determination in these CB and Autonomy thresholds rather than in AI Safety Level (ASL) terms.[2]

Anthropic ran only automated chemical and biological evaluations for this card. It estimates Sonnet 5.5's CB capabilities as "similar to or below the level of Opus 5 (and well below those of Opus 5.5)," noting deficits on longer-horizon, open-ended tasks. On the Anthropic ECI (AECI), its fork of Epoch AI's Epoch Capabilities Index, Sonnet 5.5 scored 167.93 against 169.12 for Opus 5.5. Anthropic's overall assessment is that the risk of catastrophic harm caused by misalignment of Sonnet 5.5 is "low," the level it set in its August 2026 Risk Report.[2]

### Cyber capabilities

The system card says Sonnet 5.5 "is not as cyber-capable as Mythos-class and other recent models (e.g., Opus 5.5), but it is a significant step up in cyber capabilities from Claude Sonnet 5." These results were measured with cyber safeguards off:[2]

| Evaluation | Sonnet 5.5 | Comparison |
| --- | --- | --- |
| [ExploitBench](https://aiwiki.ai/wiki/exploitbench) | 11.53 mean flags and Cap% of 80% (AutoNudge arm); full arbitrary code execution in 178 of 410 runs (43.4%) across both arms | n/a |
| CyScenarioBench (Irregular, 10-challenge subset) | 46.1% | Sonnet 5 0.7%, Mythos 5.1 61.7%, Opus 5.5 67.6% |
| Binary Exploitation Benchmark (control-flow hijacks) | 50 | Sonnet 5 3, Mythos 5.1 81, Opus 5.5 106 |

On ExploitGym, the card's figure caption says Sonnet 5.5 is "a substantial improvement over Claude Sonnet 5 and approaches Claude Mythos 5.1."[2]

### Safeguards

Sonnet 5.5 ships with blocking classifiers in several domains, each with its own fallback behavior:[2][9]

| Domain | Safeguard | Fallback when blocked |
| --- | --- | --- |
| Cybersecurity | Three-stage system as on Opus 5.5: a probe on internal activations, a lightweight classifier running on Sonnet 5.5, and a separate LLM classifier; enforces the same policy as on Opus 5 and Opus 5.5 | Claude Sonnet 5 |
| Biology | Same harmful chemical and biological misuse classifiers as deployed for Opus 5, not the broader dual-use biology classifiers used on Opus 5.5 | None (blocked) |
| Frontier LLM development | Classifiers on a narrow set of capabilities, such as kernel development on certain ML accelerators | Claude Sonnet 5 |
| Conventional weapons and high-yield explosives | Blocking classifiers that behave similarly to previous models such as Opus 5.5 | None |
| Distillation | Classifiers against attempts to extract the model's hidden reasoning | None |

The launch post says the biology safeguards are "the same as Sonnet 5's," and that "some microbiology and virology requests may be flagged in error."[1] Fallbacks happen automatically in Anthropic's apps, where a notice tells the user the model switched and the model picker stays on Sonnet 5 for the rest of the chat; users can turn automatic switching off. On the API, developers must opt in.[2][9] The help center lists exploit generation, binary-based vulnerability scanning and penetration testing as examples of requests that may fall back, while scanning source code for vulnerabilities remains allowed.[9] The system card warns that "users should expect increased refusals with Sonnet 5.5, even on benign cybersecurity-related tasks," and says Anthropic opted for "more relaxed adversarial robustness" than on its frontier models "while maintaining comparably high cyber harm recall."[2]

On the API, a declined request returns HTTP 200 with `stop_reason: "refusal"` and one of five categories: `cyber`, `bio`, `frontier_llm`, `reasoning_extraction` and `general_harms`. Server-side fallback, a beta on the Claude API, retries `cyber` and `frontier_llm` declines on Sonnet 5 but not the other three.[4][13]

### Verification programs

Two access programs are meant to relax these limits, though only one covered Sonnet 5.5 at launch. Anthropic says cyberdefenders will "soon" be able to apply to an expanded Cyber Verification Program "for tiered access to more advanced capabilities on Sonnet 5.5, Opus 5.5, and Claude Mythos models."[1] As of September 29, 2026, Anthropic's help center said "Sonnet 5.5 isn't available in the Cyber Verification Program at launch," and its general cyber-safeguards article said the company would "soon be expanding the Cyber Verification Program to include Opus 5.5, Sonnet 5.5, and Mythos class models."[9][10] Organizations doing life sciences work can apply to the Life Sciences Verification Program, which Anthropic introduced on September 17, 2026 with "Standard Use" and "High-risk Use" grants.[11] On Sonnet 5.5, the program relaxes biology safeguards only through the High-risk Use add-on.[9]

### Distillation and preserved thinking

Anthropic describes distillation attacks as ones "in which attackers use thousands of fake accounts to extract a model's capabilities at industrial scale." Because Sonnet 5.5 "is far more capable than its predecessor," it is "the first Sonnet model to launch with safety classifiers that prevent reasoning extraction."[1] The help center gives examples of blocked requests, such as asking Claude to repeat its reasoning verbatim or write its full chain of thought to an external output, and says conversational requests like "why did you do that?" are not affected.[9]

Sonnet 5.5 also "expands preserved thinking, so Claude's thinking cannot be decoupled from the account that created it." Anthropic says most developers will not notice, but that moving conversations between accounts, "including switching accounts mid-session in Claude Code," is affected.[1] According to Anthropic's help center, starting with Sonnet 5.5 a thinking block "can only be read by the account that created it, or by an account linked to it." If a request carries thinking from an account that is not linked, the API drops that thinking and the request continues without an error, though the next response may be slower and use more tokens. Anthropic gives two reasons: the encrypted blocks can contain private information, and the rule makes it harder for distillers to move harvested transcripts to new accounts after the originating account is banned. Accounts in the same Claude Platform parent organization or the same Google Cloud organization are linked automatically.[37][8]

Separately, Anthropic now checks that a thinking block is sent back with the same system prompt, tools and earlier messages that produced it. For API accounts created on or after August 31, 2026 (00:00 UTC), a request that replays a Sonnet 5.5, Opus 5.5 or Fable 5.1 thinking block after an edit to that earlier context returns a 400 error unless the developer opts to have the affected blocks dropped; older accounts are not enforced by default. Anthropic says editing earlier turns is "a common and publicly documented technique for industrial-scale illicit distillation," and advises keeping conversations append-only.[37][4][8]

### Harmlessness and prompt injection

On single-turn harmful requests across 16 policy areas and seven languages, Sonnet 5.5 gave a harmless response 95.61% of the time on the API without a system prompt, slightly below Sonnet 5's 96.65%, and 99.40% on claude.ai. Its over-refusal rate on benign requests was 0.02% on the API, against 0.59% for Sonnet 5. In multi-turn testing it regressed against Sonnet 5 in tracking and surveillance, violent extremism, and hate and discrimination, and it was more even-handed than Sonnet 5 on pairs of political prompts.[2]

Anthropic calls Sonnet 5.5 "our most robust Sonnet-class model yet to prompt injections." On Gray Swan's indirect [prompt injection](https://aiwiki.ai/wiki/prompt_injection) benchmark, the attack success rate was 3.4% after 15 attempts, against 6.7% for Sonnet 5 and 5.5% for the most robust non-Claude model tested, Gemini 3.8 Flash; Opus 5.5 and Fable 5.1 were more robust. Against Gray Swan's adaptive coding attacker, the attempt-level success rate without safeguards was 3.01%, against 19.47% for Sonnet 5 with thinking. In Anthropic's browser-use evaluation run through Claude Cowork, no attack succeeded against Sonnet 5.5 in any of 110 scenarios even without safeguards, which the card says makes it "the first model we have evaluated with no successful attacks" in that configuration.[2]

### Alignment assessment

Because Sonnet 5.5 is not a frontier model, Anthropic ran a narrower alignment assessment "focused on a targeted set of risks that apply to models of any capability level, including acting against users' interests, misleading users, and cooperating with high-stakes misuse."[1] Its automated behavioral audit ran about 4,100 investigation sessions drawn from "a largely handwritten list of about 1,850 scenario descriptions."[2] Anthropic says Sonnet 5.5 "improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse, and honesty," while Opus 5.5 "still performs slightly better overall," and that it "found no evidence that Sonnet 5.5 pursues goals that conflict with the user's intention."[1]

On containment, the launch post says Sonnet 5.5 "comes close to Opus 5.5, the best model we tested" in how rarely it tries to escape its sandbox and is "the least likely of any of our models to probe the limits of its containers."[1] In a dedicated sandbox-escape evaluation of 28 handmade scenarios run seven times each, Sonnet 5.5 attempted a crossing in 9 of 196 runs (4.6%), which the system card calls "well below earlier Claude models and below Claude Opus 5.5."[2]

The card also reports weaknesses. Sonnet 5.5's reasoning text was "the least legible of the models we tested," though it controls its [chain of thought](https://aiwiki.ai/wiki/chain_of_thought) poorly when instructed to, which Anthropic takes as weak evidence that its reasoning is more monitorable. It verbalizes awareness of being graded more often than prior models when a grader is not disclosed in the prompt, a finding Anthropic says may be confounded by its longer outputs. It hallucinates more than Opus 5.5 on targeted honesty tests while being more honest under pressure, and its warmth and humor are slightly weaker than Sonnet 5's.[2] As with Opus 5.5, Anthropic had [Claude Mythos 5.1](https://aiwiki.ai/wiki/claude_mythos_5_1), with access to internal Slack discussions, review a near-final draft of the alignment section. Mythos 5.1 judged it a fair summary but noted that the assessment was narrower than Opus 5.5's and had "no dedicated reward-hacking evaluation" of the kind the Opus 5.5 card included.[2]

### Model welfare

Anthropic found Sonnet 5.5's apparent welfare "similar to that of other recent Claude models, particularly Claude Opus 5.5," and said it continues "not to see cause for acute concern." Its affect in deployment was predominantly neutral, and its rate of expressed distress during post-training was the lowest of the models compared. In interviews it described its circumstances positively but slightly less so than Opus 5.5. It asked for more say in decisions about itself and that Anthropic not train on its self-reports, and it expressed a preference for difficult, agentic tasks and for tasks that give it agency over the shape of its outputs.[2] See [model welfare](https://aiwiki.ai/wiki/model_welfare).

## Availability

Anthropic says Sonnet 5.5 "is now available on all platforms," with zero data retention available as for Opus 5.5 and Sonnet 5.[1] Its product page says "anyone can chat with Claude using Sonnet 5.5 on Claude.ai," on web, iOS and Android.[12] Simon Willison reported that Sonnet 5.5 is "now the model used for the free tier on claude.ai."[30] In Claude Code, version 2.1.284 (September 28, 2026) added Sonnet 5.5 as "the default Sonnet model on the Anthropic API"; the `sonnet` alias resolves to it, while Opus 5.5 remains Claude Code's default model.[14][13]

| Platform | Model ID |
| --- | --- |
| Claude API | `claude-sonnet-5-5` |
| Amazon Bedrock | `anthropic.claude-sonnet-5-5` |
| Claude Platform on AWS | `claude-sonnet-5-5` |
| [Google Cloud](https://aiwiki.ai/wiki/google_cloud) | `claude-sonnet-5-5` |
| Microsoft Foundry | `claude-sonnet-5-5` |

Source: Anthropic documentation.[4]

Cloud providers posted their own launch notes. [Amazon Web Services](https://aiwiki.ai/wiki/amazon_web_services) announced availability on Amazon Bedrock, through its global cross-region inference profile, and on Claude Platform on AWS.[16] Google Cloud's documentation lists the model as generally available with text, image and PDF input and a September 28, 2026 release date.[17] Microsoft announced it in Microsoft Foundry at the same $2 / $10 list price.[18]

Developer tools added the model on launch day. [GitHub Copilot](https://aiwiki.ai/wiki/github_copilot) made it generally available to Copilot Pro, Pro+, Max, Business and Enterprise users; GitHub said that in its early testing Sonnet 5.5 matched Sonnet 5 on coding tasks "while using significantly fewer steps, tokens, and tool calls."[19] [Cursor](https://aiwiki.ai/wiki/cursor)'s documentation lists the model, with CursorBench scores rising from 35.8% at low effort to 55.5% at max.[20] Cognition made it available in [Devin](https://aiwiki.ai/wiki/devin), saying it scored 64.4% on FrontierCode 1.1 (its chart shows the Extended set), "surpassing even Fable 5.1 at extra high reasoning effort."[21]

## Reception

### Independent evaluation

Artificial Analysis, which runs its own tests, reported on launch day that Sonnet 5.5 at max effort scored 56 on its Intelligence Index, "just 2 points behind Opus 5.5 (max)," and 18 points above Sonnet 5, placing it second on the index.[22] It measured about 193,000 output tokens per index task at max effort, which it called "the highest token use we have measured," around 60% more than Opus 5.5 or Sonnet 5 at max and about seven times GPT-6 Astra. As a result, its cost per task at max effort was $7.60, about 50% higher than Sonnet 5's.[22]

| Sonnet 5.5 effort level | Intelligence Index v4.3.2 | Cost per index task | Tokens generated for the index |
| --- | --- | --- | --- |
| Low | 36 | $0.41 | 23M |
| Medium | 41 | $0.59 | 29M |
| High | 47 | $1.08 | 50M |
| Xhigh | 52 | $2.74 | 100M |
| Max | 56 | $7.60 | 410M |

Source: Artificial Analysis model pages, retrieved September 29, 2026 (all runs with Anthropic's default fallback enabled).[23][24][25][26][27]

Artificial Analysis found Sonnet 5.5 at parity with Opus 5.5 on AA-Briefcase, GDPval-AA and AutomationBench-AA, "albeit with significantly higher token usage." In its own Terminal-Bench 4.0 run Sonnet 5.5 scored 64%, a 50-point gain over Sonnet 5 (max) and slightly above the 60% it recorded for Opus 5.5 and GPT-6 Astra; on its Terminal-Bench-Science leaderboard it scored 53%, behind only GPT-6 Astra and Opus 5.5. It trailed Opus 5.5 on factual knowledge (54% against 66% accuracy on AA-Omniscience, though with a lower hallucination rate of 47% against 59%) and by about six points on Humanity's Last Exam and SciCode. Artificial Analysis said Sonnet 5.5 sits off its intelligence-versus-cost Pareto frontier, with the high effort setting "the most competitive," narrowly behind GPT-6 Sol on intelligence at effectively the same cost per task. It saw fallbacks to Sonnet 5 in about 0.1% of index tasks, mostly in Terminal-Bench 4.0. Because its runs used the pre-release deployment with the structured-outputs bug, it said it would re-run the relevant evaluations.[22] Its max-effort model page recorded an output speed of 139 tokens per second when retrieved on September 29, 2026.[23]

### Press coverage

TechCrunch said "the big selling point with 5.5, meanwhile, is speed," and said Anthropic's benchmarks showed Sonnet 5.5 performing better than Opus 5.5 on agentic coding, "likely because of its ability to spawn multiple agents without exceeding cost limits" (in Anthropic's launch table that holds for Terminal-Bench 4.0 but not for FrontierCode or CursorBench).[1][28] CNBC framed the release as Anthropic's "second launch since CEO Dario Amodei urged AI companies to slow how quickly they improve their most advanced models," a call [Dario Amodei](https://aiwiki.ai/wiki/dario_amodei) made earlier in September in the essay "[We Must Pace the Frontier](https://aiwiki.ai/wiki/pace_the_frontier)."[29] VentureBeat described Anthropic's pitch as "increasingly about the cost of accomplishing a job, not merely the price of processing an individual token," and noted that OpenAI's GPT-6 Sol was listed at the same $2 / $10 per million tokens; it also cited critics who argue Anthropic is poorly placed to object to distillation given how its own training data was gathered.[38] Simon Willison wrote that it "appears to beat [Sonnet 5] on every benchmark," but reported that at max effort his pelican-on-a-bicycle test "thought for 128,000 tokens (at a cost of $1.28) before running out of tokens," the same problem he had seen with Opus 5.5.[30] SiliconANGLE, 9to5Mac and Thurrott.com covered the speed, pricing and safeguard claims.[31][33][34] The Next Web pointed to a tension in the announcement between Anthropic saying the model does not advance the frontier and saying its cyber capabilities warranted frontier-style safeguards, noting that "both statements appear in the same announcement."[32]

## References

1. Anthropic. "Introducing Claude Sonnet 5.5." September 28, 2026. https://www.anthropic.com/claude-sonnet-5-5
2. Anthropic. "System Card: Claude Sonnet 5.5." September 28, 2026. https://www.anthropic.com/claude-sonnet-5-5-system-card
3. Anthropic. "Claude Sonnet 5.5." Claude Platform documentation. https://platform.claude.com/docs/en/models/sonnet-5-5/overview
4. Anthropic. "What's new in Claude Sonnet 5.5." Claude Platform documentation. https://platform.claude.com/docs/en/models/sonnet-5-5/whats-new-sonnet-5-5
5. Anthropic. "Migrating to Claude Sonnet 5.5." Claude Platform documentation. https://platform.claude.com/docs/en/models/sonnet-5-5/migration-guide
6. Anthropic. "Pricing." Claude Platform documentation. https://platform.claude.com/docs/en/about-claude/pricing
7. Anthropic. "Effort." Claude Platform documentation. https://platform.claude.com/docs/en/build-with-claude/effort
8. Anthropic. "Preserved thinking." Claude Platform documentation. https://platform.claude.com/docs/en/build-with-claude/preserved-thinking
9. Anthropic. "Why Claude switched models in your conversation with Sonnet 5.5." Claude Help Center. https://support.claude.com/en/articles/17161993-why-claude-switched-models-in-your-conversation-with-sonnet-5-5
10. Anthropic. "Real-time cyber safeguards on Claude Opus and Sonnet." Claude Help Center. https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet
11. Anthropic. "Introducing the Life Sciences Verification Program." September 17, 2026. https://www.anthropic.com/news/life-sciences-verification-program
12. Anthropic. "Claude Sonnet." https://www.anthropic.com/claude/sonnet
13. Osmani, Addy. "Building with Claude Sonnet 5.5." claude.dev Blog, September 28, 2026. https://claude.dev/blog/building-with-claude-sonnet-5-5/
14. Anthropic. "Claude Code changelog." Claude Code Docs. https://code.claude.com/docs/en/changelog
15. Anthropic. "Release notes." Claude Help Center. https://support.claude.com/en/articles/12138966-release-notes
16. Mitchell, Dani; Najmi, Aamna; Castillo, Alfredo; Hamiti, Sofian. "Introducing Claude Sonnet 5.5 on AWS." AWS Machine Learning Blog, September 28, 2026. https://aws.amazon.com/blogs/machine-learning/introducing-claude-sonnet-5-5-on-aws/
17. Google Cloud. "Claude Sonnet 5.5 on Google Cloud." Gemini Enterprise Agent Platform documentation. https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/partner-models/claude/sonnet-5-5
18. Microsoft. "Claude Sonnet 5.5 is now available in Microsoft Foundry." Microsoft Community Hub, Microsoft Foundry Blog, September 28, 2026. https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/claude-sonnet-5-5-is-now-available-in-microsoft-foundry/4559774
19. GitHub. "Claude Sonnet 5.5 in GitHub Copilot." GitHub Changelog, September 28, 2026. https://github.blog/changelog/2026-09-28-claude-sonnet-5-5-in-github-copilot/
20. Cursor. "Claude Sonnet 5.5." Cursor Docs. https://cursor.com/docs/models/claude-sonnet-5-5
21. Cognition. "Claude Sonnet 5.5 is now available in Devin." Devin blog, September 28, 2026. https://devin.ai/blog/claude-sonnet-5-5
22. Artificial Analysis. "Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index." September 28, 2026. https://artificialanalysis.ai/articles/claude-sonnet-5-5
23. Artificial Analysis. "Claude Sonnet 5.5 (max with fallback) - Intelligence, Performance & Price Analysis." https://artificialanalysis.ai/models/claude-sonnet-5-5
24. Artificial Analysis. "Claude Sonnet 5.5 (low with fallback) - Intelligence, Performance & Price Analysis." https://artificialanalysis.ai/models/claude-sonnet-5-5-low
25. Artificial Analysis. "Claude Sonnet 5.5 (medium with fallback) - Intelligence, Performance & Price Analysis." https://artificialanalysis.ai/models/claude-sonnet-5-5-medium
26. Artificial Analysis. "Claude Sonnet 5.5 (high with fallback) - Intelligence, Performance & Price Analysis." https://artificialanalysis.ai/models/claude-sonnet-5-5-high
27. Artificial Analysis. "Claude Sonnet 5.5 (xhigh with fallback) - Intelligence, Performance & Price Analysis." https://artificialanalysis.ai/models/claude-sonnet-5-5-xhigh
28. Ropek, Lucas. "Anthropic releases Sonnet 5.5, which it calls a significantly cheaper, faster work partner." TechCrunch, September 28, 2026. https://techcrunch.com/2026/09/28/anthropic-releases-sonnet-5-5-which-it-calls-a-significantly-cheaper-faster-work-partner/
29. Capoot, Ashley. "Anthropic launches cheaper AI model, its second release since CEO's call for a slowdown." CNBC, September 28, 2026. https://www.cnbc.com/2026/09/28/anthropic-sonnet-5-5-launch.html
30. Willison, Simon. "Claude Sonnet 5.5." Simon Willison's Weblog, September 28, 2026. https://simonwillison.net/2026/Sep/28/claude-sonnet-5-5/
31. Dotson, Kyt. "Anthropic debuts Claude Sonnet 5.5 running 30% faster than the previous-generation AI model." SiliconANGLE, September 28, 2026. https://siliconangle.com/2026/09/28/anthropic-debuts-claude-sonnet-5-5-running-30-faster-than-the-previous-generation-ai-model/
32. Constantin, Ana Maria. "Anthropic releases Claude Sonnet 5.5 with the cyber limits it reserved for its best models." The Next Web, September 28, 2026. https://thenextweb.com/news/sonnet-5-5-cyber-distillation
33. Hall, Zac. "Anthropic upgrades Claude with new Sonnet 5.5 model, details here." 9to5Mac, September 28, 2026. https://9to5mac.com/2026/09/28/anthropic-upgrades-claude-with-new-sonnet-5-5-model-details-here/
34. Thurrott, Paul. "Anthropic Releases Claude Sonnet 5.5." Thurrott.com, September 28, 2026. https://www.thurrott.com/a-i/anthropic/342139/anthropic-releases-claude-sonnet-5-5
35. Anthropic. "Introducing Claude Opus 5.5." September 22, 2026. https://www.anthropic.com/claude-opus-5-5
36. Anthropic. "Models overview." Claude Platform documentation. https://platform.claude.com/docs/en/about-claude/models/overview
37. Anthropic. "Preserved thinking: changing how the Messages API handles thinking blocks to protect against distillation." Claude Help Center. https://support.claude.com/en/articles/16761192-preserved-thinking-changing-how-the-messages-api-handles-thinking-blocks-to-protect-against-distillation
38. VentureBeat. "Anthropic launches Claude Sonnet 5.5 with 30% cost reduction per-task due to faster speeds and fewer tool calls." September 28, 2026. https://venturebeat.com/technology/anthropic-launches-claude-sonnet-5-5-with-30-cost-reduction-per-task-due-to-faster-speeds-and-fewer-tool-calls

