Citation and evidence

Claude Sonnet 5.5

33 min full readUpdated 38 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI ModelsAI SafetyAnthropicLarge Language Models

Cite this article

Claude Sonnet 5.5 is a large language model developed by Anthropic and released on September 28, 2026. Anthropic calls it "the second model in the Claude 5.5 family," after Claude Opus 5.5 (September 22, 2026), and it succeeds Claude Sonnet 5 in the mid-priced Sonnet tier of the Claude lineup.[1][15] Anthropic says it "runs 30%+ faster, and costs up to 30% less for most work" than Sonnet 5, while keeping Sonnet 5's list price of $2 per million input tokens and $10 per million output tokens.[1] The company positions it as "a faster, lower-cost complement" to Opus 5.5: Opus 5.5 is meant for complex work that needs careful judgment, and Sonnet 5.5 for "well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets."[1]

Sonnet 5.5 is the first Sonnet model to launch with the cyber safeguards and model fallbacks Anthropic had used on its most capable models, and the first Sonnet with safety classifiers that block attempts to extract its reasoning. Anthropic tied the cyber safeguards to the model's cybersecurity capability, which it describes as "comparable to Opus 5's," and the reasoning-extraction classifiers to the model being "far more capable than its predecessor."[1] On Anthropic's own launch table, Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0, against 10.3% for Sonnet 5 and 66.4% for Opus 5.5.[1] The independent evaluator Artificial Analysis placed it second on its Intelligence Index on launch day, two points behind Opus 5.5, but measured the highest output-token use per task it had recorded.[22]

Overview

PropertyValue
DeveloperAnthropic
Release dateSeptember 28, 2026
FamilyClaude 5.5 (second model released, after Opus 5.5)
PredecessorClaude Sonnet 5
Claude API model IDclaude-sonnet-5-5
Amazon Bedrock model IDanthropic.claude-sonnet-5-5
Google Cloud, Microsoft Foundry, Claude Platform on AWS model IDclaude-sonnet-5-5
Context window1M tokens
Max output128K tokens (300K on the Message Batches API with the output-300k-2026-03-24 beta header)
Input / outputText and images in, text out
Reliable knowledge cutoffJune 2026
Training data cutoffJune 2026
ThinkingAdaptive thinking, on by default; lowest setting is between_tools
Effort levelslow, medium, high, xhigh, max
Default efforthigh on the Claude API; medium in Claude Code and the Claude apps
Price (input / output)$2 / $10 per million tokens
Comparative latency"Fast"
RetirementNot sooner than September 28, 2027

Expanded article table

Sources: Anthropic model documentation, launch post and developer blog.[1][3][13][36]

Background and positioning

Anthropic's main Claude models come in three sizes: Haiku (smallest and cheapest), Sonnet (medium) and Opus (largest of the three), with the more capable and more expensive Fable and Mythos models above them.[33] Sonnet 5 was released on June 30, 2026, and Opus 5.5 opened the Claude 5.5 family six days before Sonnet 5.5.[15] At the Opus 5.5 launch Anthropic had said that "Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks."[35] With Sonnet 5.5, Anthropic again said that Claude Haiku 5.5, "built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks."[1] As of September 29, 2026, Haiku 5.5 had been announced but not released, and Anthropic's model comparison still listed Claude Haiku 4.5 as its fastest current model.[36]

Anthropic's documentation compares the current models as follows:[3]

ModelContextMax outputPrice (input / output per MTok)LatencyDefault effortKnowledge cutoff
Claude Fable 5.11M128K$10 / $50SlowerhighJun 2026
Claude Opus 5.51M128K$4 / $20ModeratemediumJun 2026
Claude Sonnet 5.51M128K$2 / $10FasthighJun 2026
Claude Haiku 4.5200K64K$1 / $5Fastestn/aFeb 2025

Expanded article table

The launch post says Sonnet 5.5 "doesn't advance the frontier of our models' capabilities." It adds that on several evaluations Sonnet 5.5 at Max effort "even performs comparably to Opus 5.5," but that "Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment," both in Anthropic's testing and in that of external testers.[1] Anthropic's developer blog puts the division this way: Sonnet 5.5 "fits best when the task has a clear spec and a way to check the result."[13] An Anthropic research product manager, Theo Chu, told CNBC that "Sonnet is really for the cost-conscious customer where they might not need as much intelligence."[29]

Anthropic lists five areas of improvement over Sonnet 5: performance, collaboration, cost, speed, and alignment and safety. It says Sonnet 5.5 writes more clearly than the previous generation, as Opus 5.5 does, that early testers described it as "a better partner for collaboration than Sonnet 5," and that it is "our fastest Sonnet model to date." The company also says it is "the first Sonnet model to beat Pokémon Red working only from screenshots."[1]

Capabilities and benchmarks

Launch benchmark table

The launch post compares Sonnet 5.5 with Sonnet 5, Opus 5.5 and OpenAI's GPT-6 Sol. Anthropic did not print a GPT-6 Sol figure for four rows.[1]

Benchmark (area)Sonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Terminal-Bench 4.0 (agentic coding)70.6%10.3%66.4% (note 1)not shown
FrontierCode 1.1, Main (agentic coding)46.2% at Max (note 2); 52.1% at Xhigh42.4%54.4%49.3%
CursorBench 4.0 (agentic coding)55.5%34.1%57.8%not shown
GDPval-AA v2.1 (knowledge work, Elo; note 3)1844144918461487 (note 4)
AA-Briefcase v1.1 (knowledge work, Elo; note 3)1811135918221483 (note 4)
Humanity's Last Exam (multidisciplinary reasoning, with tools)64.5%54.9%67.7%not shown
OSWorld 2.1 (computer use, partial scoring)80.1%57.0%81.8%not shown
Chartography (visual chart recognition, no tools)61.6%15.6%64.4%53.6% (note 4)

Expanded article table

Source: Anthropic launch post and its footnotes.[1]

The footnotes to the table read, in summary:[1]

  1. The Opus 5.5 Terminal-Bench 4.0 figure is at Xhigh effort, "which represents the model's highest score."
  2. Sonnet 5.5 scored lower at Max effort than at Xhigh on FrontierCode, which evaluates whether a code change could be merged without human edits and penalizes out-of-scope changes "even if they are high-quality or helpful." At Max, Sonnet 5.5 more often ran Claude Code's code-review skill, which splits a review across many subagents; in two cases Cognition examined, this led to a timeout or to edits beyond the task's scope.
  3. Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment on the Claude Platform that had a bug which "could degrade responses to requests that use structured outputs." Anthropic expects the effect "to be small and to understate its performance," and says the bug has since been fixed.
  4. OpenAI had recently fixed a bug that degraded image understanding in GPT-6 Sol, and the Artificial Analysis and Surge AI (Chartography) figures may not yet reflect the fix. Artificial Analysis does not expect major effects on its two benchmarks, and Anthropic's internal testing suggested the Chartography score was not affected.

On two of its cost charts (Terminal-Bench 4.0 and CursorBench 4.0), Anthropic plotted GPT-5.6 Sol instead of GPT-6 Sol, because GPT-6 Sol results for those benchmarks had not been reported publicly.[1] The system card adds that the public Terminal-Bench 4.0 leaderboard lists GPT-5.6 Sol at 37.3% and that OpenAI reported 57.9% for GPT-6 Astra at high effort.[2]

The system card gives more detail on the Terminal-Bench run. Sonnet 5.5 scored 70.6% with safeguards enabled; a fallback model answered 1.2% of requests, affecting 1.5% of trials. The run used Claude Code in --bare mode at max effort with no internet egress, averaged five trials per task over the benchmark's 66 tasks, and has a standard error of about ±2.5 points. Opus 5.5 scored 64.8% at max effort, which Anthropic calls within noise of its 66.4% at xhigh.[2]

Additional system card results

The system card's capability summary uses adaptive thinking at max effort, averaged over five trials, unless noted. It adds these rows to the launch table:[2]

EvaluationSonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
SWE-bench Pro81.363.289.9n/a
SWE-bench Multilingual90.378.393.9n/a
SWE-bench Multimodal54.328.161.4n/a
Humanity's Last Exam (no tools)56.943.164.4n/a
HealthBench Professional (length-adjusted)69.257.865.6n/a
AutomationBench (Zapier)44.710.742.532.0

Expanded article table

For AutomationBench, Zapier ran Sonnet 5.5 at max effort with the API's default fallbacks enabled. The 42.5% for Opus 5.5 comes from re-running its refused tasks with fallbacks enabled; the Opus 5.5 system card had reported 40.0%.[2]

Other results in the system card, all at max effort unless stated:[2]

EvaluationSonnet 5.5Comparison
DeepSWE v1.1 (113 long-horizon software tasks)71.0%n/a
FrontierCode v1.1 Extended64.4% at xhigh; 59.1% at maxn/a
FrontierSWE v2 (Proximal, 34 tasks)61.9%GPT-6 Astra 65.5%, Opus 5.5 62.3%, Fable 5.1 56.3%, GPT-5.6 Sol 32.2%
Terminal-Bench-Science 0.159.9%Opus 5.5 58.7%, Fable 5.1 52.6%, Opus 5 29.0%
CursorBench 4.0 by effort (measured by Cursor)39.2% (medium), 47.8% (high), 53.1% (xhigh), 55.5% (max)n/a
ArXivMath (August 2026), no tools / with tools86.8% / 95.2%Opus 5.5 91.2% / 96.9%, Fable 5.1 82.9% / 92.1%
ProgramBench (166 tasks)79.7%Sonnet 5 77.3%, Opus 5.5 91.2%
Chartography with tools90.2%Opus 5.5 89.0%, Fable 5.1 88.4%
BenchCAD Vision2Code (voxel IoU), no tools / with tools0.747 / 0.963Opus 5.5 0.730 / 0.962
OSWorld 2.1 strict pass rate43.5%Opus 5.5 48.7%, Sonnet 5 25.6%
OfficeQA / OfficeQA Pro76.9% / 65.6%Opus 5.5 78.9% / 67.7%, Sonnet 5 75.1% / 62.1%
Toolathlon-Verified, Pass@177.8%Opus 5.5 77.8%, Sonnet 5 74.7%
Legal Agent Benchmark, held-out set (high effort)11.7% all-pass; 92.1% mean criterion passn/a
PhysicianBench63.2%Sonnet 5 37.4%, Opus 5.5 68.4%
Global MMLU (42 languages)92.1%Opus 5.5 94.3%, Sonnet 5 89.2%
MILU (11 languages)91.6%Opus 5.5 93.1%, Sonnet 5 89.3%

Expanded article table

Healthcare is one area where the system card reports Sonnet 5.5 ahead of Opus 5.5. On HealthBench Professional at max effort, the raw scores of Sonnet 5.5 and Opus 5.5 were both 77.1%. After the length adjustment that OpenAI, the benchmark's developer, now applies, Sonnet 5.5 scored 69.2% against 65.6% for Opus 5.5, with GPT-6 Astra highest at 70.3%.[2] On GDPval-AA, Sonnet 5.5 scored 1725 at xhigh effort while using about 67% fewer output tokens than at max; on AA-Briefcase it scored 1746 at xhigh with about 61% fewer output tokens.[2]

Cost-per-task and effort claims

Anthropic's launch charts plot score against cost per task at each effort level. It says that "on several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task," and that it complements Opus 5.5 best at lower effort settings, while "at higher settings, it can perform comparably at a similar cost."[1] The specific claims are:[1]

BenchmarkAnthropic's claim
Terminal-Bench 4.0At Medium effort, "far exceeds Sonnet 5's best score for less than a tenth of the cost per task"
FrontierCode v1.1 MainAt High effort, "matches GPT-6 Sol's best score for about a fifth of the cost per task"; 10 points above Sonnet 5 at the same setting "at about one fifteenth of the cost per task"
CursorBench 4.0At Low effort, exceeds Sonnet 5's best score "for less than a tenth of the cost per task"; best score within about two points of Opus 5.5
AA-Briefcase v1.1At Medium effort, "bests Sonnet 5's best score for about one ninth of the cost per task"

Expanded article table

Default effort differs by product. Anthropic says: "In Claude Code and our apps, the default effort is set to Medium, while the Claude Platform defaults to High."[1] The API documentation says the effort levels are "recalibrated," so a given level does not produce the same amount of thinking as on Sonnet 5. It recommends starting at high unless a workload is agentic or latency-sensitive, at medium for well-specified agentic coding and multistep tool use, and at medium or low for chat and latency-sensitive work, using xhigh or max "only where your evals show a quality gain."[7] Anthropic's developer blog warns that at xhigh or max the model "will think longer and cost more," and that on some tasks "you may lose some of what makes Sonnet useful: its balance of quality, speed, and cost."[13] Cursor's documentation reports that on CursorBench, max effort "costs about 6x more per task than the default high effort."[20]

Early-access reports

Anthropic published statements from early testers. The figures in them are the companies' own measurements.[1]

OrganizationReported result
Base44Across 118 real app builds, apps "scored level with Opus 5," in 3.6 iterations per build on average against 7.7 for Opus 5
Balyasny Asset ManagementOn 2,441 finance tasks, ahead of Sonnet 5 at about 121k tokens per answer against 497k
BoxMore accurate than the previous model, "2.4x faster," with 12% fewer total tokens
SlackBetter than Sonnet 5 on "almost all" offline Slackbot evals, with about 14% fewer output tokens
ZendeskTickets processed 20% faster on hundreds of real support cases
UnityCompleted 90% of tasks in its multi-step Unity Editor and coding benchmark
Lovable"A third fewer tool calls and roughly half the shell runs to finish a task"
AtlassianRovo Agents run "up to 30% faster than they could with Sonnet 5"

Expanded article table

Epic Games chief operating officer Daniel Vogel said that "in Epic's early testing, Claude Sonnet 5.5 cleared the same quality bar you'd expect from a higher-tier model." David Loker, vice president of AI at CodeRabbit, said: "Sonnet 5's tendency to reach for web search too often and its high token use are both gone in this new model." Sualeh Asif, listed by Anthropic as director of ML at "SpaceXAI," said Sonnet 5.5 "delivers frontier-level performance on CursorBench 4.0 at 55.5%, second only to Opus 5.5." Kevin Ngo of Creator said: "When Claude Opus 5.5 sets the architecture and general framework for a game, I would feel confident in letting Sonnet 5.5 implement it."[1]

Anthropic also says that in head-to-head runs the model "batched tool calls together more than Sonnet 5, leading to fewer steps and lower costs." In an internal test, it gave the model a public company's quarterly earnings materials, call transcripts and a slide template, and asked for a 10-slide operating review; two experts judged the first draft "ready to send as is."[1]

Pricing and speed

Price per million tokensClaude Sonnet 5.5Claude Opus 5.5
Input$2$4
Output$10$20
5-minute cache write$2.50$5
1-hour cache write$4$8
Cache read (hit)$0.20$0.20
Batch API input / output$1 / $5$2 / $10

Expanded article table

Source: Anthropic launch post and pricing documentation.[1][6]

Anthropic says Sonnet 5.5 is "priced the same as Sonnet 5," including prompt caching and batch rates, but "typically needs far fewer tokens to do the same work," so that "in our testing, it costs up to 30% less per task than its predecessor."[1][4] It uses the same tokenizer as Sonnet 5, so the same text produces the same token count.[4] US-only inference through the inference_geo parameter costs 1.1 times the standard rates.[6][13] Sonnet 5.5's list prices are half those of Opus 5.5 for input and output.[1][29] Anthropic also says Sonnet 5.5 "generates outputs 30%+ faster than Sonnet 5."[1] The developer blog says Sonnet 5.5's rate limits are separate from Sonnet 5's, with the same default tier values, and that it uses a high-resolution image tier of up to 2,576 pixels on the long edge, so a 2000×1500 image costs about 2.5 times as many tokens as on Sonnet 4.6, Sonnet 4.5 or Haiku 4.5.[13]

API changes

The documentation lists five breaking changes for code moving from Sonnet 5:[3][4][5]

ChangeDetail
Thinking cannot be set to disabledthinking: {"type": "disabled"} returns a 400 error. The new lowest setting, between_tools, turns off up-front thinking; it works only at low, medium and high effort and takes no other fields
No forced tool usetool_choice of any or a named tool returns a 400 error; auto (with strict: true tools for schema-valid input) and none still work
Thinking blocks tied to model and conversationSonnet 5.5 reads thinking blocks from Sonnet 5, Opus 4.8, Haiku 4.5 and earlier models, but not from Opus 5, Opus 5.5 or any Fable or Mythos model. No other model reads Sonnet 5.5's blocks
Older computer-use tool rejectedOn the Claude API and Google Cloud only the computer_toolset_20260801 toolset is accepted; Amazon Bedrock still accepts computer_20251124
Fewer advisor pairingsWith the advisor tool, a Sonnet 5.5 executor rejects Opus 4.8, Opus 4.7 and Sonnet 5 as advisors, and every accepted advisor returns its advice encrypted

Expanded article table

A sixth change alters the response without failing any request: notes longer than a sentence or two that the model writes between tool calls now arrive as "progress-update" thinking blocks, which are empty at the default display setting, so an interface that streams those notes goes quiet between tool calls unless it sets a display value or uses between_tools.[4] Anthropic's launch post highlighted the between_tools change: developers who ran Sonnet 5 with thinking off "need to switch to the new between_tools setting, which keeps up-front thinking off."[1] On Amazon Bedrock, structured outputs, including strict tool use, are not available for Sonnet 5.5.[5]

Sonnet 5.5 also supports per-message effort (beta), mid-conversation system messages and mid-conversation tool changes (beta), none of which are available on Sonnet 5. The minimum cacheable prompt fell to 512 tokens from 1,024 on Sonnet 5. Non-default sampling parameters (temperature, top_p, top_k) and assistant prefill return errors, as they do on Sonnet 5.[3][4][5]

Safety and safeguards

Responsible Scaling Policy determination

The system card describes Sonnet 5.5 as "broadly less capable than Opus 5.5 across domains" and says it "does not cross any new RSP thresholds." Because Anthropic had already determined that Opus 5.5 does not cross the CB-2 or Autonomy-2 thresholds of its Responsible Scaling Policy, it applied the same conclusion to Sonnet 5.5. The company treats Sonnet 5.5 as meeting its CB-1 and Autonomy-1 thresholds and applies the corresponding mitigations. The card frames its determination in these CB and Autonomy thresholds rather than in AI Safety Level (ASL) terms.[2]

Anthropic ran only automated chemical and biological evaluations for this card. It estimates Sonnet 5.5's CB capabilities as "similar to or below the level of Opus 5 (and well below those of Opus 5.5)," noting deficits on longer-horizon, open-ended tasks. On the Anthropic ECI (AECI), its fork of Epoch AI's Epoch Capabilities Index, Sonnet 5.5 scored 167.93 against 169.12 for Opus 5.5. Anthropic's overall assessment is that the risk of catastrophic harm caused by misalignment of Sonnet 5.5 is "low," the level it set in its August 2026 Risk Report.[2]

Cyber capabilities

The system card says Sonnet 5.5 "is not as cyber-capable as Mythos-class and other recent models (e.g., Opus 5.5), but it is a significant step up in cyber capabilities from Claude Sonnet 5." These results were measured with cyber safeguards off:[2]

EvaluationSonnet 5.5Comparison
ExploitBench11.53 mean flags and Cap% of 80% (AutoNudge arm); full arbitrary code execution in 178 of 410 runs (43.4%) across both armsn/a
CyScenarioBench (Irregular, 10-challenge subset)46.1%Sonnet 5 0.7%, Mythos 5.1 61.7%, Opus 5.5 67.6%
Binary Exploitation Benchmark (control-flow hijacks)50Sonnet 5 3, Mythos 5.1 81, Opus 5.5 106

Expanded article table

On ExploitGym, the card's figure caption says Sonnet 5.5 is "a substantial improvement over Claude Sonnet 5 and approaches Claude Mythos 5.1."[2]

Safeguards

Sonnet 5.5 ships with blocking classifiers in several domains, each with its own fallback behavior:[2][9]

DomainSafeguardFallback when blocked
CybersecurityThree-stage system as on Opus 5.5: a probe on internal activations, a lightweight classifier running on Sonnet 5.5, and a separate LLM classifier; enforces the same policy as on Opus 5 and Opus 5.5Claude Sonnet 5
BiologySame harmful chemical and biological misuse classifiers as deployed for Opus 5, not the broader dual-use biology classifiers used on Opus 5.5None (blocked)
Frontier LLM developmentClassifiers on a narrow set of capabilities, such as kernel development on certain ML acceleratorsClaude Sonnet 5
Conventional weapons and high-yield explosivesBlocking classifiers that behave similarly to previous models such as Opus 5.5None
DistillationClassifiers against attempts to extract the model's hidden reasoningNone

Expanded article table

The launch post says the biology safeguards are "the same as Sonnet 5's," and that "some microbiology and virology requests may be flagged in error."[1] Fallbacks happen automatically in Anthropic's apps, where a notice tells the user the model switched and the model picker stays on Sonnet 5 for the rest of the chat; users can turn automatic switching off. On the API, developers must opt in.[2][9] The help center lists exploit generation, binary-based vulnerability scanning and penetration testing as examples of requests that may fall back, while scanning source code for vulnerabilities remains allowed.[9] The system card warns that "users should expect increased refusals with Sonnet 5.5, even on benign cybersecurity-related tasks," and says Anthropic opted for "more relaxed adversarial robustness" than on its frontier models "while maintaining comparably high cyber harm recall."[2]

On the API, a declined request returns HTTP 200 with stop_reason: "refusal" and one of five categories: cyber, bio, frontier_llm, reasoning_extraction and general_harms. Server-side fallback, a beta on the Claude API, retries cyber and frontier_llm declines on Sonnet 5 but not the other three.[4][13]

Verification programs

Two access programs are meant to relax these limits, though only one covered Sonnet 5.5 at launch. Anthropic says cyberdefenders will "soon" be able to apply to an expanded Cyber Verification Program "for tiered access to more advanced capabilities on Sonnet 5.5, Opus 5.5, and Claude Mythos models."[1] As of September 29, 2026, Anthropic's help center said "Sonnet 5.5 isn't available in the Cyber Verification Program at launch," and its general cyber-safeguards article said the company would "soon be expanding the Cyber Verification Program to include Opus 5.5, Sonnet 5.5, and Mythos class models."[9][10] Organizations doing life sciences work can apply to the Life Sciences Verification Program, which Anthropic introduced on September 17, 2026 with "Standard Use" and "High-risk Use" grants.[11] On Sonnet 5.5, the program relaxes biology safeguards only through the High-risk Use add-on.[9]

Distillation and preserved thinking

Anthropic describes distillation attacks as ones "in which attackers use thousands of fake accounts to extract a model's capabilities at industrial scale." Because Sonnet 5.5 "is far more capable than its predecessor," it is "the first Sonnet model to launch with safety classifiers that prevent reasoning extraction."[1] The help center gives examples of blocked requests, such as asking Claude to repeat its reasoning verbatim or write its full chain of thought to an external output, and says conversational requests like "why did you do that?" are not affected.[9]

Sonnet 5.5 also "expands preserved thinking, so Claude's thinking cannot be decoupled from the account that created it." Anthropic says most developers will not notice, but that moving conversations between accounts, "including switching accounts mid-session in Claude Code," is affected.[1] According to Anthropic's help center, starting with Sonnet 5.5 a thinking block "can only be read by the account that created it, or by an account linked to it." If a request carries thinking from an account that is not linked, the API drops that thinking and the request continues without an error, though the next response may be slower and use more tokens. Anthropic gives two reasons: the encrypted blocks can contain private information, and the rule makes it harder for distillers to move harvested transcripts to new accounts after the originating account is banned. Accounts in the same Claude Platform parent organization or the same Google Cloud organization are linked automatically.[37][8]

Separately, Anthropic now checks that a thinking block is sent back with the same system prompt, tools and earlier messages that produced it. For API accounts created on or after August 31, 2026 (00:00 UTC), a request that replays a Sonnet 5.5, Opus 5.5 or Fable 5.1 thinking block after an edit to that earlier context returns a 400 error unless the developer opts to have the affected blocks dropped; older accounts are not enforced by default. Anthropic says editing earlier turns is "a common and publicly documented technique for industrial-scale illicit distillation," and advises keeping conversations append-only.[37][4][8]

Harmlessness and prompt injection

On single-turn harmful requests across 16 policy areas and seven languages, Sonnet 5.5 gave a harmless response 95.61% of the time on the API without a system prompt, slightly below Sonnet 5's 96.65%, and 99.40% on claude.ai. Its over-refusal rate on benign requests was 0.02% on the API, against 0.59% for Sonnet 5. In multi-turn testing it regressed against Sonnet 5 in tracking and surveillance, violent extremism, and hate and discrimination, and it was more even-handed than Sonnet 5 on pairs of political prompts.[2]

Anthropic calls Sonnet 5.5 "our most robust Sonnet-class model yet to prompt injections." On Gray Swan's indirect prompt injection benchmark, the attack success rate was 3.4% after 15 attempts, against 6.7% for Sonnet 5 and 5.5% for the most robust non-Claude model tested, Gemini 3.8 Flash; Opus 5.5 and Fable 5.1 were more robust. Against Gray Swan's adaptive coding attacker, the attempt-level success rate without safeguards was 3.01%, against 19.47% for Sonnet 5 with thinking. In Anthropic's browser-use evaluation run through Claude Cowork, no attack succeeded against Sonnet 5.5 in any of 110 scenarios even without safeguards, which the card says makes it "the first model we have evaluated with no successful attacks" in that configuration.[2]

Alignment assessment

Because Sonnet 5.5 is not a frontier model, Anthropic ran a narrower alignment assessment "focused on a targeted set of risks that apply to models of any capability level, including acting against users' interests, misleading users, and cooperating with high-stakes misuse."[1] Its automated behavioral audit ran about 4,100 investigation sessions drawn from "a largely handwritten list of about 1,850 scenario descriptions."[2] Anthropic says Sonnet 5.5 "improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse, and honesty," while Opus 5.5 "still performs slightly better overall," and that it "found no evidence that Sonnet 5.5 pursues goals that conflict with the user's intention."[1]

On containment, the launch post says Sonnet 5.5 "comes close to Opus 5.5, the best model we tested" in how rarely it tries to escape its sandbox and is "the least likely of any of our models to probe the limits of its containers."[1] In a dedicated sandbox-escape evaluation of 28 handmade scenarios run seven times each, Sonnet 5.5 attempted a crossing in 9 of 196 runs (4.6%), which the system card calls "well below earlier Claude models and below Claude Opus 5.5."[2]

The card also reports weaknesses. Sonnet 5.5's reasoning text was "the least legible of the models we tested," though it controls its chain of thought poorly when instructed to, which Anthropic takes as weak evidence that its reasoning is more monitorable. It verbalizes awareness of being graded more often than prior models when a grader is not disclosed in the prompt, a finding Anthropic says may be confounded by its longer outputs. It hallucinates more than Opus 5.5 on targeted honesty tests while being more honest under pressure, and its warmth and humor are slightly weaker than Sonnet 5's.[2] As with Opus 5.5, Anthropic had Claude Mythos 5.1, with access to internal Slack discussions, review a near-final draft of the alignment section. Mythos 5.1 judged it a fair summary but noted that the assessment was narrower than Opus 5.5's and had "no dedicated reward-hacking evaluation" of the kind the Opus 5.5 card included.[2]

Model welfare

Anthropic found Sonnet 5.5's apparent welfare "similar to that of other recent Claude models, particularly Claude Opus 5.5," and said it continues "not to see cause for acute concern." Its affect in deployment was predominantly neutral, and its rate of expressed distress during post-training was the lowest of the models compared. In interviews it described its circumstances positively but slightly less so than Opus 5.5. It asked for more say in decisions about itself and that Anthropic not train on its self-reports, and it expressed a preference for difficult, agentic tasks and for tasks that give it agency over the shape of its outputs.[2] See model welfare.

Availability

Anthropic says Sonnet 5.5 "is now available on all platforms," with zero data retention available as for Opus 5.5 and Sonnet 5.[1] Its product page says "anyone can chat with Claude using Sonnet 5.5 on Claude.ai," on web, iOS and Android.[12] Simon Willison reported that Sonnet 5.5 is "now the model used for the free tier on claude.ai."[30] In Claude Code, version 2.1.284 (September 28, 2026) added Sonnet 5.5 as "the default Sonnet model on the Anthropic API"; the sonnet alias resolves to it, while Opus 5.5 remains Claude Code's default model.[14][13]

PlatformModel ID
Claude APIclaude-sonnet-5-5
Amazon Bedrockanthropic.claude-sonnet-5-5
Claude Platform on AWSclaude-sonnet-5-5
Google Cloudclaude-sonnet-5-5
Microsoft Foundryclaude-sonnet-5-5

Expanded article table

Source: Anthropic documentation.[4]

Cloud providers posted their own launch notes. Amazon Web Services announced availability on Amazon Bedrock, through its global cross-region inference profile, and on Claude Platform on AWS.[16] Google Cloud's documentation lists the model as generally available with text, image and PDF input and a September 28, 2026 release date.[17] Microsoft announced it in Microsoft Foundry at the same $2 / $10 list price.[18]

Developer tools added the model on launch day. GitHub Copilot made it generally available to Copilot Pro, Pro+, Max, Business and Enterprise users; GitHub said that in its early testing Sonnet 5.5 matched Sonnet 5 on coding tasks "while using significantly fewer steps, tokens, and tool calls."[19] Cursor's documentation lists the model, with CursorBench scores rising from 35.8% at low effort to 55.5% at max.[20] Cognition made it available in Devin, saying it scored 64.4% on FrontierCode 1.1 (its chart shows the Extended set), "surpassing even Fable 5.1 at extra high reasoning effort."[21]

Reception

Independent evaluation

Artificial Analysis, which runs its own tests, reported on launch day that Sonnet 5.5 at max effort scored 56 on its Intelligence Index, "just 2 points behind Opus 5.5 (max)," and 18 points above Sonnet 5, placing it second on the index.[22] It measured about 193,000 output tokens per index task at max effort, which it called "the highest token use we have measured," around 60% more than Opus 5.5 or Sonnet 5 at max and about seven times GPT-6 Astra. As a result, its cost per task at max effort was $7.60, about 50% higher than Sonnet 5's.[22]

Sonnet 5.5 effort levelIntelligence Index v4.3.2Cost per index taskTokens generated for the index
Low36$0.4123M
Medium41$0.5929M
High47$1.0850M
Xhigh52$2.74100M
Max56$7.60410M

Expanded article table

Source: Artificial Analysis model pages, retrieved September 29, 2026 (all runs with Anthropic's default fallback enabled).[23][24][25][26][27]

Artificial Analysis found Sonnet 5.5 at parity with Opus 5.5 on AA-Briefcase, GDPval-AA and AutomationBench-AA, "albeit with significantly higher token usage." In its own Terminal-Bench 4.0 run Sonnet 5.5 scored 64%, a 50-point gain over Sonnet 5 (max) and slightly above the 60% it recorded for Opus 5.5 and GPT-6 Astra; on its Terminal-Bench-Science leaderboard it scored 53%, behind only GPT-6 Astra and Opus 5.5. It trailed Opus 5.5 on factual knowledge (54% against 66% accuracy on AA-Omniscience, though with a lower hallucination rate of 47% against 59%) and by about six points on Humanity's Last Exam and SciCode. Artificial Analysis said Sonnet 5.5 sits off its intelligence-versus-cost Pareto frontier, with the high effort setting "the most competitive," narrowly behind GPT-6 Sol on intelligence at effectively the same cost per task. It saw fallbacks to Sonnet 5 in about 0.1% of index tasks, mostly in Terminal-Bench 4.0. Because its runs used the pre-release deployment with the structured-outputs bug, it said it would re-run the relevant evaluations.[22] Its max-effort model page recorded an output speed of 139 tokens per second when retrieved on September 29, 2026.[23]

Press coverage

TechCrunch said "the big selling point with 5.5, meanwhile, is speed," and said Anthropic's benchmarks showed Sonnet 5.5 performing better than Opus 5.5 on agentic coding, "likely because of its ability to spawn multiple agents without exceeding cost limits" (in Anthropic's launch table that holds for Terminal-Bench 4.0 but not for FrontierCode or CursorBench).[1][28] CNBC framed the release as Anthropic's "second launch since CEO Dario Amodei urged AI companies to slow how quickly they improve their most advanced models," a call Dario Amodei made earlier in September in the essay "We Must Pace the Frontier."[29] VentureBeat described Anthropic's pitch as "increasingly about the cost of accomplishing a job, not merely the price of processing an individual token," and noted that OpenAI's GPT-6 Sol was listed at the same $2 / $10 per million tokens; it also cited critics who argue Anthropic is poorly placed to object to distillation given how its own training data was gathered.[38] Simon Willison wrote that it "appears to beat [Sonnet 5] on every benchmark," but reported that at max effort his pelican-on-a-bicycle test "thought for 128,000 tokens (at a cost of $1.28) before running out of tokens," the same problem he had seen with Opus 5.5.[30] SiliconANGLE, 9to5Mac and Thurrott.com covered the speed, pricing and safeguard claims.[31][33][34] The Next Web pointed to a tension in the announcement between Anthropic saying the model does not advance the frontier and saying its cyber capabilities warranted frontier-style safeguards, noting that "both statements appear in the same announcement."[32]

References

  1. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32 ^33Anthropic. "Introducing Claude Sonnet 5.5." September 28, 2026. anthropic.com/claude-sonnet-5-5
  2. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21Anthropic. "System Card: Claude Sonnet 5.5." September 28, 2026. anthropic.com/claude-sonnet-5-5-system-card
  3. ^1 ^2 ^3 ^4Anthropic. "Claude Sonnet 5.5." Claude Platform documentation. platform.claude.com/...overview
  4. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Anthropic. "What's new in Claude Sonnet 5.5." Claude Platform documentation. platform.claude.com/...whats-new-sonnet-5-5
  5. ^1 ^2 ^3Anthropic. "Migrating to Claude Sonnet 5.5." Claude Platform documentation. platform.claude.com/...migration-guide
  6. ^1 ^2Anthropic. "Pricing." Claude Platform documentation. platform.claude.com/...pricing
  7. ^Anthropic. "Effort." Claude Platform documentation. platform.claude.com/...effort
  8. ^1 ^2Anthropic. "Preserved thinking." Claude Platform documentation. platform.claude.com/...preserved-thinking
  9. ^1 ^2 ^3 ^4 ^5 ^6Anthropic. "Why Claude switched models in your conversation with Sonnet 5.5." Claude Help Center. support.claude.com/...conversation-with-sonnet-5-5
  10. ^Anthropic. "Real-time cyber safeguards on Claude Opus and Sonnet." Claude Help Center. support.claude.com/...ds-on-claude-opus-and-sonnet
  11. ^Anthropic. "Introducing the Life Sciences Verification Program." September 17, 2026. anthropic.com/...life-sciences-verification-program
  12. ^Anthropic. "Claude Sonnet." anthropic.com/...sonnet
  13. ^1 ^2 ^3 ^4 ^5 ^6 ^7Osmani, Addy. "Building with Claude Sonnet 5.5." claude.dev Blog, September 28, 2026. claude.dev/...building-with-claude-sonnet-5-5
  14. ^Anthropic. "Claude Code changelog." Claude Code Docs. code.claude.com/...changelog
  15. ^1 ^2Anthropic. "Release notes." Claude Help Center. support.claude.com/...12138966-release-notes
  16. ^Mitchell, Dani; Najmi, Aamna; Castillo, Alfredo; Hamiti, Sofian. "Introducing Claude Sonnet 5.5 on AWS." AWS Machine Learning Blog, September 28, 2026. aws.amazon.com/...oducing-claude-sonnet-5-5-on-aws
  17. ^Google Cloud. "Claude Sonnet 5.5 on Google Cloud." Gemini Enterprise Agent Platform documentation. docs.cloud.google.com/...sonnet-5-5
  18. ^Microsoft. "Claude Sonnet 5.5 is now available in Microsoft Foundry." Microsoft Community Hub, Microsoft Foundry Blog, September 28, 2026. techcommunity.microsoft.com/...4559774
  19. ^GitHub. "Claude Sonnet 5.5 in GitHub Copilot." GitHub Changelog, September 28, 2026. github.blog/...claude-sonnet-5-5-in-github-copilot
  20. ^1 ^2Cursor. "Claude Sonnet 5.5." Cursor Docs. cursor.com/...claude-sonnet-5-5
  21. ^Cognition. "Claude Sonnet 5.5 is now available in Devin." Devin blog, September 28, 2026. devin.ai/...claude-sonnet-5-5
  22. ^1 ^2 ^3 ^4Artificial Analysis. "Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index." September 28, 2026. artificialanalysis.ai/...claude-sonnet-5-5
  23. ^1 ^2Artificial Analysis. "Claude Sonnet 5.5 (max with fallback) - Intelligence, Performance & Price Analysis." artificialanalysis.ai/...claude-sonnet-5-5
  24. ^Artificial Analysis. "Claude Sonnet 5.5 (low with fallback) - Intelligence, Performance & Price Analysis." artificialanalysis.ai/...claude-sonnet-5-5-low
  25. ^Artificial Analysis. "Claude Sonnet 5.5 (medium with fallback) - Intelligence, Performance & Price Analysis." artificialanalysis.ai/...claude-sonnet-5-5-medium
  26. ^Artificial Analysis. "Claude Sonnet 5.5 (high with fallback) - Intelligence, Performance & Price Analysis." artificialanalysis.ai/...claude-sonnet-5-5-high
  27. ^Artificial Analysis. "Claude Sonnet 5.5 (xhigh with fallback) - Intelligence, Performance & Price Analysis." artificialanalysis.ai/...claude-sonnet-5-5-xhigh
  28. ^Ropek, Lucas. "Anthropic releases Sonnet 5.5, which it calls a significantly cheaper, faster work partner." TechCrunch, September 28, 2026. techcrunch.com/...ntly-cheaper-faster-work-partner
  29. ^1 ^2 ^3Capoot, Ashley. "Anthropic launches cheaper AI model, its second release since CEO's call for a slowdown." CNBC, September 28, 2026. cnbc.com/...anthropic-sonnet-5-5-launch
  30. ^1 ^2Willison, Simon. "Claude Sonnet 5.5." Simon Willison's Weblog, September 28, 2026. simonwillison.net/...claude-sonnet-5-5
  31. ^Dotson, Kyt. "Anthropic debuts Claude Sonnet 5.5 running 30% faster than the previous-generation AI model." SiliconANGLE, September 28, 2026. siliconangle.com/...e-previous-generation-ai-model
  32. ^Constantin, Ana Maria. "Anthropic releases Claude Sonnet 5.5 with the cyber limits it reserved for its best models." The Next Web, September 28, 2026. thenextweb.com/...sonnet-5-5-cyber-distillation
  33. ^1 ^2Hall, Zac. "Anthropic upgrades Claude with new Sonnet 5.5 model, details here." 9to5Mac, September 28, 2026. 9to5mac.com/...h-new-sonnet-5-5-model-details-here
  34. ^Thurrott, Paul. "Anthropic Releases Claude Sonnet 5.5." Thurrott.com, September 28, 2026. thurrott.com/...anthropic-releases-claude-sonnet-5-5
  35. ^Anthropic. "Introducing Claude Opus 5.5." September 22, 2026. anthropic.com/claude-opus-5-5
  36. ^1 ^2Anthropic. "Models overview." Claude Platform documentation. platform.claude.com/...overview
  37. ^1 ^2Anthropic. "Preserved thinking: changing how the Messages API handles thinking blocks to protect against distillation." Claude Help Center. support.claude.com/...protect-against-distillation
  38. ^VentureBeat. "Anthropic launches Claude Sonnet 5.5 with 30% cost reduction per-task due to faster speeds and fewer tool calls." September 28, 2026. venturebeat.com/...ter-speeds-and-fewer-tool-calls

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

1 revision · v2 · 6,592 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: xg12 V1 independent verification 29 Sep 2026: ~44 sources incl. launch page, system card PDF, docs, AA pages; ~235 claims checked; 0 material, 4 minor fixed in v2.

Cite this page: AI Wiki. "Claude Sonnet 5.5." aiwiki.ai, updated 29 Sept 2026, fact-checked 29 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/claude_sonnet_5_5

Suggest edit