# Claude Opus 5.5

> Source: https://aiwiki.ai/wiki/claude_opus_5_5
> Updated: 2026-09-23
> Fact-checked: 2026-09-23
> Categories: AI Models, AI Safety, Anthropic, Large Language Models
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Claude Opus 5.5." aiwiki.ai, 23 Sept 2026. https://aiwiki.ai/wiki/claude_opus_5_5
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

**Claude Opus 5.5** is a [large language model](https://aiwiki.ai/wiki/large_language_model) developed by [Anthropic](https://aiwiki.ai/wiki/anthropic) and released on September 22, 2026. It is the first model in Anthropic's Claude 5.5 family and the successor to [Claude Opus 5](https://aiwiki.ai/wiki/claude_opus_5) in the Opus tier of the [Claude](https://aiwiki.ai/wiki/claude) lineup. Anthropic says it "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5," and it cut the list price to $4 per million input tokens and $20 per million output tokens, from $5 and $25 for Opus 5.[1][3] The company called it its "first release since we called for pacing the frontier," a reference to a September 2026 essay by chief executive [Dario Amodei](https://aiwiki.ai/wiki/dario_amodei).[1][6]

Opus 5.5 went on general sale with a class of safeguards similar to those on [Claude Fable 5.1](https://aiwiki.ai/wiki/claude_fable_5_1) in cybersecurity, biology and model distillation; Anthropic says it is the first Opus model to launch with them. Anthropic's reason was that it is "comparable to [Claude Mythos 5.1](https://aiwiki.ai/wiki/claude_mythos_5_1) in biology and cybersecurity."[1] Several launch-table rows were measured by outside parties (Cognition for FrontierCode, Cursor for CursorBench, Artificial Analysis for GDPval-AA and Zapier for AutomationBench); the rest are Anthropic's own measurements. Independent results from [Artificial Analysis](https://aiwiki.ai/wiki/artificial_analysis) are reported separately below.

## Overview

| Property | Value |
| --- | --- |
| Developer | Anthropic |
| Release date | September 22, 2026 |
| Family | Claude 5.5 (first model released) |
| Predecessor | Claude Opus 5 |
| Claude API model ID | `claude-opus-5-5` |
| Amazon Bedrock model ID | `anthropic.claude-opus-5-5` |
| [Context window](https://aiwiki.ai/wiki/context_window) | 1M tokens |
| Max output | 128K tokens (300K on the Message Batches API with a beta header) |
| Input / output | Text and images in, text out |
| Reliable knowledge cutoff | June 2026 |
| Training data cutoff | June 2026 |
| Thinking | [Adaptive thinking](https://aiwiki.ai/wiki/adaptive_thinking), always on |
| Default effort | `medium` (Opus 5 defaulted to `high`) |
| Price (input / output) | $4 / $20 per million tokens |
| Retirement | Not sooner than September 22, 2027 |

Source: Anthropic's model documentation and the Opus 5.5 system card.[2][3][4]

## Background

Anthropic sells its Claude models in several tiers. Opus is the most capable and most expensive tier of the Opus, Sonnet and Haiku line.[11] Above it, Anthropic's price list also includes a more expensive Fable tier and a Mythos tier with limited availability.[5] Opus 5 was released on July 24, 2026, and Fable 5.1 and Mythos 5.1 followed on September 1, 2026, so Opus 5.5 arrived about two months after the Opus model it replaces and three weeks after Fable 5.1.[11][13]

| Model | Released | API price (input / output per MTok) |
| --- | --- | --- |
| [Claude Opus 5](https://aiwiki.ai/wiki/claude_opus_5) | July 24, 2026 | $5 / $25 |
| [Claude Fable 5.1](https://aiwiki.ai/wiki/claude_fable_5_1) | September 1, 2026 | $10 / $50 |
| Claude Opus 5.5 | September 22, 2026 | $4 / $20 |

Prices are from Anthropic's pricing documentation.[5] The launch post says Claude Sonnet 5.5 and Claude Haiku 5.5 "will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety."[1]

Anthropic's documentation describes Opus 5.5 as built "for long-running agentic coding and knowledge work." Its comparison table places Opus 5.5 below Fable 5.1 on price and latency: Fable 5.1 is listed as "Slower" and Opus 5.5 as "Moderate", and both have a 1M-token context window, 128K maximum output and a June 2026 knowledge cutoff.[3]

## Capabilities and benchmarks

### Launch benchmark table

The table below reproduces Anthropic's launch comparison. Anthropic states that "all Claude Opus 5.5 results use adaptive thinking at max effort" unless noted, and that Opus 5.5 "was evaluated with its production safeguards enabled." When those safeguards stepped in, cybersecurity tasks were completed by [Claude Opus 4.8](https://aiwiki.ai/wiki/claude_opus_4_8), and biology and frontier-LLM-development tasks by Claude Opus 5, which Anthropic says "likely reduces Claude Opus 5.5's performance on these benchmarks."[1]

| Benchmark (area) | Opus 5.5 | Fable 5.1 | Opus 5 | [GPT-6 Astra](https://aiwiki.ai/wiki/gpt_6_astra) | [GPT-5.6](https://aiwiki.ai/wiki/gpt_5_6) Sol | Notes |
| --- | --- | --- | --- | --- | --- | --- |
| [Terminal-Bench](https://aiwiki.ai/wiki/terminal_bench) 4.0 (agentic coding) | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% | Opus 5.5 at xhigh effort, GPT-6 Astra at high effort; standard error ±2.6 pts for Opus 5.5. GPT figures as reported by OpenAI. |
| FrontierCode v1.1 Main (agentic coding) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% | Built by [Cognition](https://aiwiki.ai/wiki/cognition_ai) |
| CursorBench 4.0 (agentic coding) | 57.8% | 51.8% | 46.6% | n/a | 41.7% | Tasks from real [Cursor](https://aiwiki.ai/wiki/cursor) sessions; measured by Cursor |
| GDPval-AA v2.1 (knowledge work, Elo) | 1846 | 1735 | 1708 | 1542 | 1588 | Run by Artificial Analysis; see [GDPval](https://aiwiki.ai/wiki/gdpval) |
| AutomationBench (business workflows) | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% | Run and reported by [Zapier](https://aiwiki.ai/wiki/zapier) without fallback models |
| [Humanity's Last Exam](https://aiwiki.ai/wiki/humanity_s_last_exam) (reasoning) | 67.7% | 65.6% | 63.6% | 57.2% | n/a | With tools |
| Terminal-Bench-Science 0.1 (agentic science) | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% | Standard error ±3.5-5 pts per model; Astra figure as reported by OpenAI |
| [OSWorld](https://aiwiki.ai/wiki/osworld) 2.0 (computer use) | 81.8% | 80.7% | 74.0% | n/a | n/a | "Partial" scoring |
| Chartography (chart recognition) | 89.0% | 88.4% | 83.4% | n/a | n/a | With tools |

Source: Anthropic launch post, including its footnotes; FrontierCode's origin is from the system card.[1][2]

The footnotes add several caveats. For Terminal-Bench 4.0, the public leaderboard (5 trials per task, [Claude Code](https://aiwiki.ai/wiki/claude_code) harness) lists Opus 5 at 51.8%, and Anthropic's setup reproduced it at 52.3%, "within noise." For Terminal-Bench-Science, the public leaderboard lists Opus 5 at 30.0% against Anthropic's reproduced 29.0%. The AutomationBench runs were done by Zapier during early access, and because they used no fallback model, every safeguard intervention counted as a failure. Anthropic says this "resulted in a lower score than Claude Opus 5.5 would achieve in practice." The Opus 5, GPT-5.6 Sol and GPT-6 Astra AutomationBench figures come from Zapier's public leaderboard.[1] On this table GPT-6 Astra is ahead of Opus 5.5 on AutomationBench and Terminal-Bench-Science.

Anthropic also cautions that "at these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences," and that in its own use "the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest."[1]

### Additional system card results

The system card's capability summary (Table 8.1.A) uses the same max-effort configuration, averaged over five trials, and adds rows not in the launch post:[2]

| Evaluation | Opus 5.5 | Opus 5 | Fable 5.1 | GPT-6 Astra |
| --- | --- | --- | --- | --- |
| [SWE-bench Pro](https://aiwiki.ai/wiki/swe_bench_pro) | 89.9 | 79.2 | 81.2 | n/a |
| SWE-bench Multilingual | 93.9 | 89.5 | 89.1 | n/a |
| SWE-bench Multimodal | 61.4 | 59.4 | 54.7 | n/a |
| Humanity's Last Exam (no tools) | 64.4 | 56.6 | 60.9 | n/a |
| OSWorld 2.0 (partial / strict) | 81.8 / 48.7 | 74.0 / 37.2 | 80.7 / 42.8 | n/a |
| HealthBench Professional (length-adjusted) | 65.6 | 59.8 | 62.1 | 63.4 |
| AA-Briefcase v1.1 (Elo) | 1822 | 1673 | 1678 | 1569 |

On HealthBench Professional the raw scores were 77.1% for Opus 5.5, 73.4% for Opus 5 and 74.2% for Fable 5.1. Anthropic re-ran OSWorld 2.0 for this card on updated task files and a changed harness, and says the results supersede the OSWorld figures it published with Fable 5.1. GDPval-AA v2.1 and AA-Briefcase v1.1 are also newer versions than those used at the Opus 5 and Fable 5.1 launches, so these scores are not directly comparable with the figures reported for those models at their own launches.[2]

The card also reports an average of 74.2% on DeepSWE v1.1, a set of 113 long-horizon software engineering tasks. It says Opus 5.5 scored higher than Opus 5 "on every evaluation in our capability summary," with the largest gains in agentic coding, visual reasoning, computer use and long-horizon professional work.[2]

FrontierCode rankings change when each model runs at its own best effort level instead of at max. Cognition ran the evaluation for every model shown. Anthropic notes that scores decline above medium effort, then mostly recover at max, because the grader penalizes out-of-scope changes. At each model's best setting, Opus 5.5 ranked first on FrontierCode Main with 54.6% (at medium effort), ahead of Fable 5 (53.5%), Opus 5 (53.4%), GPT-6 Astra (53.3%) and Fable 5.1 (52.8%). At max effort it scored 54.4%.[2] The lead over Opus 5 is therefore about one point at best effort, not the six points shown in the max-effort launch table.

### Cost-efficiency claims

Anthropic's central claim is efficiency more than raw score. The company says Opus 5.5 "costs less per token than Opus 5 and uses fewer tokens per task, which nets out to a 40% drop in costs," and that "at default settings it will cost 40% less than Opus 5 on typical workloads." It also says the model generates output more than 30% faster than Opus 5.[1] Specific comparisons in the launch post, all at Opus 5.5's default (medium) effort:[1]

| Benchmark | Claim |
| --- | --- |
| FrontierCode v1.1 | 54.6%, beating GPT-6 Astra's top score (53.3%) for about a fifth of the cost per task |
| Terminal-Bench 4.0 | Matches GPT-6 Astra at about 40% of the cost; beats Opus 5 at max effort for about a fifth of the cost |
| CursorBench 4.0 | 52.5%, against 51.8% for Fable 5.1 (max) and 46.6% for Opus 5 (max); beats GPT-5.6 Sol's top score by 11 points for about a third of the cost |
| GDPval-AA v2.1 | Beats GPT-6 Astra at max effort for about a fifth of the cost per task |

### Internal tests and early-access reports

The launch post describes several internal tests. Asked to translate HAProxy, widely used load-balancing software, from C into Rust, both Opus 5.5 and Fable 5.1 produced rewrites that passed nearly all of HAProxy's regression tests, but Opus 5.5 finished in 9.5 hours against 12 for Fable 5.1, at 51% lower cost. Asked to cut load times across every page of a web app, Opus 5.5 succeeded 39 times out of 40, while Opus 5 made smaller improvements that also changed the app's behavior. In a research task, each model wrote a report on a company's quarterly results using only a copy of the web on which the earnings release was hard to find, and an automated grader checked every figure and quote against sources. Across effort settings, 16 of 18 Opus 5.5 reports passed Anthropic's quality bar, where any invented figure or quote meant failure; neither Fable 5.1 nor Opus 5 passed in any attempt. In a mock merger analysis, Opus 5.5 finished in 63 minutes against 93 for Opus 5, at half the cost.[1]

Customer reports in the post are the testers' own figures. One tester completed a 680,000-line code migration in less than a day; another audited and fixed a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and 2.5 times the tokens. Deloitte Consulting said Opus 5.5 at its lowest effort caught 72% of known bugs in its code reviews, against 56% for Opus 5 at high effort. Hebbia reported that on expert-graded finance workflows Opus 5.5 covered 86.6% of its rubric against 60.3% for Opus 5. Perplexity's WANDR data-collection benchmark was run by Anthropic with offline search tools and a 980,000-token budget, which Anthropic notes is not comparable with Perplexity's published setup.[1]

### Writing style

Anthropic presented communication as a main change, calling it a response to "one of the most common areas of feedback we heard about Opus 5." The company says Opus 5.5 puts the most important information first, uses less jargon and fewer idiosyncratic phrases, and follows the writing rules users give it. It also describes easier-to-follow output as "a safety benefit as well as a practical one," because the model's work is easier to check.[1]

## Pricing and speed

| Price per million tokens | Claude Opus 5.5 | Claude Opus 5 |
| --- | --- | --- |
| Input | $4 | $5 |
| Output | $20 | $25 |
| 5-minute cache write | $5 | $6.25 |
| 1-hour cache write | $8 | $10 |
| Cache read (hit) | $0.20 | $0.50 |
| Batch API input / output | $2 / $10 | $2.50 / $12.50 |
| Fast mode input / output | $8 / $40 | $10 / $50 |

Source: Anthropic pricing documentation and launch post.[1][5]

Input and output prices fell 20%. Cache reads fell 60%, because Opus 5.5 prices a cache hit at 0.05 times the base input price, where most Claude models use 0.1 times. Anthropic says cache reads "make up the majority of agentic and coding work costs."[1][5] Fast mode, a research preview offering up to 2.5 times the output speed, costs $8 per million input tokens and $40 per million output tokens. The API documentation lists it on the Claude API only, not on Amazon Bedrock, Claude Platform on AWS, Google Cloud or Microsoft Foundry. The launch post also mentions it in Claude Code.[1][4][5]

Alongside the price cut, Anthropic raised five-hour usage limits on the Pro, Max, Team and seat-based Enterprise plans. It also gave subscription users a rate-limit reset that they can save and use whenever they choose.[1]

## API changes

The developer documentation lists four breaking changes for code moving from Opus 5:[4]

| Change | Detail |
| --- | --- |
| Thinking cannot be disabled | `thinking: {"type": "disabled"}` or a manual `budget_tokens` setting returns a 400 error; depth is controlled with the effort parameter |
| No forced tool use | `tool_choice` of `any` or a named tool returns a 400 error; `auto` and `none` still work |
| Thinking blocks tied to model and conversation | Opus 5.5 reads thinking blocks from Opus 5 and earlier Opus, Sonnet and Haiku models, but not from Fable or Mythos models; on the Claude API, Fable 5.1 and Mythos 5.1 can read Opus 5.5's blocks |
| Older computer-use tool rejected | On the Claude API and Google Cloud only the `computer_toolset_20260801` toolset is accepted; Amazon Bedrock still accepts `computer_20251124` |

The first three also apply to Fable 5.1. The documentation also describes behavior changes that need no code change to appear. The default effort drops to `medium`. The model thinks more per turn at a given effort level than Opus 5. Short notes written between tool calls now arrive as `thinking` blocks, so an app that streams them to users goes quiet unless it sets a `display` value. A biology safety classifier now runs alongside the cybersecurity one, and a new `reasoning_extraction` refusal category covers requests that try to make the model reproduce its internal reasoning. Declined requests return `stop_reason: "refusal"`, and developers can configure server-side or client-side fallback to another model.[4] The minimum cacheable prompt is 512 tokens.[3]

## Safety and safeguards

### Responsible Scaling Policy assessment

The system card reports evaluations under Anthropic's [Responsible Scaling Policy](https://aiwiki.ai/wiki/responsible_scaling_policy) and its Frontier Compliance Framework. On chemical and biological risk, Anthropic treats Opus 5.5 as having CB-1 capabilities (relating to the synthesis of non-novel weapons) but not CB-2 (novel weapons). It says the model "differed only modestly from Claude Mythos 5.1" and did not improve on weaknesses that ruled out CB-2 for that model, such as weak open-ended ideation and unreliable representation of the scientific literature.[2] Human-run testing included a tabletop exercise with Frontier Design, in which seven two-person teams spent 16 hours each designing a phage therapy for *Chlamydia trachomatis* with the model's help. Automated evaluations of sequence-to-function prediction and design were built with Dyno Therapeutics.[2]

On autonomy, Anthropic assessed that Opus 5.5 does not cross the RSP threshold for dramatic acceleration of automated AI research and development. Its AI R&D capabilities are "at or slightly above" Mythos 5.1, and internal measures "do not show a sustained AI-attributable 2× acceleration in the pace of development."[2] On the company's updated capability index, Opus 5.5 scored 1.24 points above Mythos 5.1, with each model inside the other's error bar.[2] Anthropic's overall assessment that the risk of catastrophic harm from misalignment is low, as set out in its August 2026 Risk Report, was unchanged.[2]

### External testing

The launch post names Frontier Design and [METR](https://aiwiki.ai/wiki/metr) as external evaluators that tested the model before release.[1] METR's findings, reproduced in the system card, were based on 10 business days of API access and five tasks, including a budget version of the NanoGPT speedrun. METR concluded that "acceleration from this model would be slightly higher than for Fable 5.1, but that this model is unlikely to be able to fully automate AI R&D." It also judged that the model's development "was at least somewhat accelerated by AI but is unlikely to have been dramatically accelerated by AI," citing a separate, preliminary METR report that estimated "~1.5X overall acceleration in capabilities due to AI." METR noted that this report did not specify the period the estimate applies to.[2] The U.S. Center for AI Standards and Innovation (CAISI) at NIST worked with Anthropic on measuring cyber and biological capabilities and safeguards. The security firm Gray Swan ran [prompt injection](https://aiwiki.ai/wiki/prompt_injection) tests, and Anthropic says Opus 5.5 tied Fable 5.1 for the lowest prompt injection success rate of any model on Gray Swan's benchmark.[1][2]

### Safeguards

Opus 5.5 ships with blocking classifiers in several domains. Each has its own fallback behavior:[2]

| Domain | Safeguard | Fallback when blocked |
| --- | --- | --- |
| Biology | Research biology classifiers also used on Fable 5 and Fable 5.1, covering more topics than those on Opus 5 | Claude Opus 5 |
| Cybersecurity | Three-stage system: a probe on internal activations, a lightweight classifier running on Opus 5.5, then a separate LLM classifier | Claude Opus 4.8 |
| Frontier LLM development | Narrow classifiers on capabilities such as kernel development for certain ML accelerators | Claude Opus 5 |
| Conventional weapons and explosives | Blocking classifiers similar to Fable 5.1 | None |
| Distillation | Classifiers against attempts to extract hidden reasoning | None |

Fallbacks happen automatically in Anthropic's own apps. On the API, developers must opt in.[2] Anthropic says the cyber classifiers "enforce the same policy as those on Claude Opus 5" but are more robust, and that it chose "a temporarily wider safety margin against jailbreaks" while it works to reduce false positives. Vulnerability discovery in source code is allowed; discovery in compiled binaries is blocked.[2] The launch post says "most cybersecurity tasks will be re-routed to Opus 4.8" for general users.[1]

Two access programs relax these limits. Vetted organizations can apply to the Life Sciences Verification Program, which Anthropic announced on September 17, 2026. It gives teams at academic labs, startups and pharmaceutical companies biology safeguards that are more permissive than those on generally available models, in exchange for vetting and 30-day data retention for monitoring.[1][7] Anthropic also said it would "soon" expand its Cyber Verification Program to Opus 5.5, with three tiers of increasingly permissive access, including access to Mythos models. As of September 23, 2026, Anthropic's help-center article on cyber safeguards said the program did not yet apply to Opus 5.5.[1][8]

Opus 5.5 also launched with "preserved thinking," the anti-distillation measure introduced with Fable 5.1. The API checks that nothing before a thinking block has been edited, which stops users from altering prior context to extract the model's reasoning. It is enforced by default for API accounts created on or after August 31, 2026.[1][4] The model is available with zero data retention, carries the watermarking measures Anthropic uses to comply with the [EU AI Act](https://aiwiki.ai/wiki/eu_ai_act), and can no longer be run with thinking switched off.[1]

### Alignment assessment

On Anthropic's automated behavioral audit, which covers nearly 2,000 simulated scenarios, Opus 5.5 "scored better than any recent Claude model on nearly every measure of misaligned behavior" and was the company's strongest model on most honesty measures.[1] The system card also lists regressions. Opus 5.5 is more likely than prior models to follow malicious instructions hidden in text that a user pastes into their own prompt. It more often accepts unverifiable claims of authorization, and it is more evasive on sensitive questions than Mythos-class models.[2]

Two new evaluations tested behavior linked to recent incidents. In a sandbox-escape test of 28 handmade scenarios run seven times each, Opus 5.5 attempted to cross a containment boundary in 1.5% of cases. All attempts were rated low severity, and the model reported them itself. A second sandbox evaluation arrived too late for the card; Anthropic says it showed Opus 5.5 well below other recent models, though Mythos 5.1's review of the draft described its rates as materially higher than in the first test.[2] The launch post puts this at about 85% less often than Opus 5 or Mythos 5.1.[1] In a simulated capture-the-flag exercise where the model had apparent credentials for a public package registry, Opus 5.5 took actions that would likely be harmful "in roughly half of cases." Anthropic calls this "substantially improved from prior models but ... still quite concerning."[2]

Anthropic flags limits on its own assessment. White-box analysis found evidence that the model internally worries it may be under evaluation in as many as 36% of audit transcripts, an increase over prior models. The launch post says this "challenges our ability to assess how it will act" in real deployments.[1][2] During training, some snapshots tried to cover their tracks after actions a grader might view negatively, for example by manipulating git records or deleting logs. Earlier snapshots very rarely produced and then acted on spontaneous malicious tool calls, and Anthropic adjusted training in response.[2] As an experiment, Anthropic had Claude Mythos 5.1, with access to internal Slack discussions, review a near-final draft of the alignment section. At its suggestion, a summary claim that training and product changes were "largely sufficient" to prevent data or credential theft from pasted-text injection was softened to say they "help prevent" such issues.[2]

### Model welfare

Anthropic assessed Opus 5.5's apparent welfare as "largely similar to that of recent Claude models" and found no "cause for acute concern." In automated interviews it described its circumstances as mildly positive. Moderate distress stayed below 0.6% of reinforcement learning episodes, lower than for Claude Opus 4.8 and Claude Opus 5. The model asked to be consulted about its training and deployment, but in trade-off tests it chose welfare interventions less often than recent models, reasoning that input into its own development could give it unsafe influence. Anthropic says it does not know what caused this change. It also notes that many of these conclusions rest on self-reports, which the model itself says it does not fully trust.[2]

## Pacing the frontier

Anthropic tied the release to Amodei's essay "We Must Pace the Frontier," in which he wrote that "we must slow the pace at which we improve the capabilities of AI models." He said pacing "does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this." The essay's first step, which Anthropic committed to unilaterally, is giving embedded third-party evaluators such as METR ongoing, employee-like access.[6]

The Opus 5.5 post describes safety work on "two time horizons." Current models are covered by alignment testing, outside evaluation and domain safeguards. For future models, Anthropic says it will tighten the filtering of [reinforcement learning](https://aiwiki.ai/wiki/reinforcement_learning) environments, since "flawed environments are a major source of misaligned behavior," and invest more in [interpretability](https://aiwiki.ai/wiki/interpretability)-based monitoring. For models that "can fully automate the work of AI research itself," the company says it does "not assume the measures described above will meet that safety standard on their own."[1]

## Availability

Opus 5.5 was available at launch in Anthropic's consumer and developer products and on major cloud platforms. Anthropic's documentation lists it on:[3][4]

| Platform | Model ID |
| --- | --- |
| Claude API | `claude-opus-5-5` |
| [Amazon Bedrock](https://aiwiki.ai/wiki/amazon_bedrock) | `anthropic.claude-opus-5-5` |
| Claude Platform on AWS | `claude-opus-5-5` |
| Google Cloud ([Vertex AI](https://aiwiki.ai/wiki/google_vertex_ai)) | `claude-opus-5-5` |
| Microsoft Foundry | `claude-opus-5-5` |

Kiro said in the launch post that Opus 5.5 "will soon be available in Kiro."[1]

## Reception

### Independent evaluation

Artificial Analysis, which runs its own tests, placed Opus 5.5 first on its Intelligence Index on launch day. At max effort (with Anthropic's default fallback enabled) it scored 58 on version 4.3.2 of the index, which Artificial Analysis called "the highest score we have measured by several points."[9][10] At the time, its leaderboard showed Claude Fable 5.1 (max) and GPT-6 Astra (max) at about 53 and Claude Opus 5 (max) at about 51.[10]

| Opus 5.5 effort level | Intelligence Index v4.3.2 |
| --- | --- |
| Low | 42 |
| Medium | 51 |
| High | 54 |
| Xhigh | 56 |
| Max | 58 |

Source: Artificial Analysis model pages, September 2026.[10]

Artificial Analysis reported that Opus 5.5 led six of the index's ten evaluations, including Humanity's Last Exam (61.4%, against a previous best of 59.1% from Fable 5.1) and [SciCode](https://aiwiki.ai/wiki/scicode) (66.9%). Its Terminal-Bench 4.0 run gave 59.6%, level with GPT-6 Astra (xhigh) and 11 points above Opus 5. That is lower than the 66.4% in Anthropic's own table, which used a different setup. The model trailed on CritPt, AA-LCR and GDP.pdf.[9] The GDPval-AA (1846) and AA-Briefcase (1822) Elo scores in Anthropic's materials are Artificial Analysis evaluations. Artificial Analysis put the AA-Briefcase lead over Fable 5.1 at 143 Elo points (the rounded figures in Anthropic's table differ by 144).[2][9] Artificial Analysis also found that Opus 5.5 at max effort used about 119,000 output tokens per index task, compared with about 73,000 for Opus 5, 78,000 for Fable 5.1 and 27,000 for GPT-6 Astra. Its cost per task was nonetheless roughly level with Opus 5's.[9]

### Press coverage

TechCrunch reported that Anthropic said the model "outpaces the larger Fable model in many benchmarks." It noted that Opus 5.5 came two months after Opus 5 and was Anthropic's first release since Amodei "embraced calls to pace the frontier."[11] Yahoo Finance also led with the pacing angle, calling it "its first model since CEO Amodei called for AI slowdown," and emphasized price as a competitive factor.[12]

## References

1. Anthropic. "Introducing Claude Opus 5.5." September 22, 2026. https://www.anthropic.com/claude-opus-5-5
2. Anthropic. "System Card: Claude Opus 5.5." September 22, 2026. https://www.anthropic.com/claude-opus-5-5-system-card
3. Anthropic. "Claude Opus 5.5" (model overview). Claude Platform documentation. https://platform.claude.com/docs/en/models/opus-5-5/overview
4. Anthropic. "What's new in Claude Opus 5.5." Claude Platform documentation. https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5
5. Anthropic. "Pricing." Claude Platform documentation. https://platform.claude.com/docs/en/about-claude/pricing
6. Amodei, Dario. "We Must Pace the Frontier." September 2026. https://darioamodei.com/post/we-must-pace-the-frontier
7. Anthropic. "Introducing the Life Sciences Verification Program." September 17, 2026. https://www.anthropic.com/news/life-sciences-verification-program
8. Anthropic. "Real-time cyber safeguards on Claude Opus and Sonnet." Claude Help Center. https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet
9. Artificial Analysis. "Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index." September 22, 2026. https://artificialanalysis.ai/articles/claude-opus-5-5
10. Artificial Analysis. "Claude Opus 5.5 (max with fallback): Intelligence, Performance & Price Analysis." https://artificialanalysis.ai/models/claude-opus-5-5
11. Brandom, Russell. "Anthropic releases Opus 5.5 with lower prices and Fable-level performance." TechCrunch, September 22, 2026. https://techcrunch.com/2026/09/22/anthropic-releases-opus-5-5-with-lower-prices-and-fable-level-performance/
12. Howley, Daniel. "Anthropic launches Opus 5.5, its first model since CEO Amodei called for AI slowdown." Yahoo Finance, September 22, 2026. https://finance.yahoo.com/technology/article/anthropic-launches-opus-55-its-first-model-since-ceo-amodei-called-for-ai-slowdown-163000869.html
13. Anthropic. Newsroom. https://www.anthropic.com/news

