Claude Haiku 5.5
Claude Haiku 5.5 is a large language model released by Anthropic on October 7, 2026. It belongs to the small, fast tier of the Claude family. Anthropic positions it for high-volume work such as summarization, classification, database queries, customer support, and narrowly scoped coding subagents.[1]
Its API provides a one-million-token context window, a standard output limit of 128,000 tokens, and adaptive thinking. The default effort is medium.[10][6]
Model specifications and availability
| Property | Claude Haiku 5.5 |
|---|---|
| Release date | October 7, 2026 |
| Claude API model ID | claude-haiku-5-5 |
| Input and output | Text and images to text |
| Context window | 1,000,000 tokens |
| Standard maximum output | 128,000 tokens |
| Reliable knowledge cutoff | June 2026 |
| Training data cutoff | June 2026 |
| Default effort | medium |
| Retirement commitment | Not before October 7, 2027 |
These are the published API specifications.[2]
| Platform | Published model ID |
|---|---|
| Claude API | claude-haiku-5-5 |
| Amazon Bedrock | anthropic.claude-haiku-5-5 |
| Claude Platform on AWS | claude-haiku-5-5 |
| Google Cloud | claude-haiku-5-5 |
| Microsoft Foundry | claude-haiku-5-5 |
The platform IDs above are listed in the model overview.[2]
The Claude API ID is fixed, without a date suffix or separate alias. Haiku 4.5 integrations need their platform's replacement ID.[7]
Pricing and prompt length
Haiku 5.5 has two prompt-length pricing tiers. The following Claude API list prices are in US dollars per million tokens, as listed on October 8, 2026. A prompt longer than 100,000 tokens uses the higher tier, including the higher output rate.[1]
Prompt caching
Prompt caching reuses an unchanged prefix containing tool definitions, a system prompt, and conversation content. A cache hit reduces the price of those input tokens; it does not remove them from the request. The response distinguishes uncached input_tokens, cache_creation_input_tokens, and cache_read_input_tokens.[3]
For Haiku 5.5, the published minimum cacheable prefix is 512 tokens on the Claude API, Claude Platform on AWS, Google Cloud, and Microsoft Foundry. A shorter marked prefix runs without caching rather than producing a caching error. Bedrock has its own caching documentation and usage fields. Five-minute entries refresh when reused; a one-hour duration is available by setting ttl: "1h" in cache_control.[3]
Changing top-level effort between requests invalidates cached conversation messages. Haiku 5.5 supports a beta per-message effort change on the Claude API and Google Cloud that preserves the earlier cache. It requires adaptive thinking and the mid-conversation-output-config-2026-07-01 beta header; it cannot change effort while thinking is disabled.[6]
Batch processing and extended output
The Message Batches API processes requests asynchronously at half the standard input and output rates. It is suited to work that does not require an immediate answer. Each batch is limited to 100,000 requests or 256 MB, whichever comes first, and unfinished requests expire after 24 hours.[4]
For Haiku 5.5, the output-300k-2026-03-24 beta header raises the batch output limit to 300,000 tokens. This is not the synchronous Messages API's standard 128,000-token limit. The extended-output beta is available on the Claude API and Claude Platform on AWS, not Amazon Bedrock, Google Cloud, or Microsoft Foundry. A single large generation can take more than an hour.[4]
Thinking and effort
Adaptive thinking is enabled by default. The model decides how much reasoning to allocate to a request and can reason between tool calls. An effort level is a behavioral control, not a guaranteed token budget. max_tokens is the hard per-request limit on thinking plus the answer, so a small value can leave no room for visible text.[5]
Thinking tokens are billed as output even when their text is omitted from the response. usage.output_tokens_details.thinking_tokens reports the raw reasoning-token count; output_tokens remains the inclusive billed total. A tool-use loop makes multiple requests, each with its own output cap, so one request's max_tokens does not cap the cost of the entire task.[5]
| Effort | Documented starting use |
|---|---|
low | Short tool tasks, chat, and simple high-volume requests |
medium | Default; most work, including agentic coding |
high | Knowledge work, longer tasks, and stricter instruction following |
xhigh | Work where application evaluations justify additional reasoning |
max | Work where application evaluations justify the highest token expenditure |
Anthropic recommends comparing higher-effort Haiku configurations with Claude Sonnet 5.5 on quality, latency, and cost. Haiku permits thinking: {"type": "disabled"} at low, medium, or high; disabling thinking at xhigh or max returns a 400 error.[6]
The response can begin with a thinking block rather than answer text. Thinking text is omitted by default. To receive summarized reasoning, set thinking.display to "summarized". Applications displaying the answer should select text blocks by type instead of assuming the first block is text.[10]
Context and token accounting
The context window covers the system prompt, tool definitions, messages, tool results, image and document tokens, and the output being generated. Cached prefixes still occupy it. Haiku 5.5 preserves earlier thinking blocks by default, unlike Haiku 4.5, which strips previous thinking blocks from subsequent context.[8]
The one-million-token window is the default and does not need a long-context beta header. Requests can contain up to 600 images or PDF pages, but request-size limits can be reached before the token limit. Input that already exceeds the window returns a 400 error. If generation reaches the window limit, the response can stop with stop_reason: "model_context_window_exceeded".[8]
Haiku 5.5 does not receive the injected remaining-context-budget tags used by Haiku 4.5. Developers can supply an explicit task budget through the task-budgets beta instead.[8]
Haiku 5.5 uses the newer Claude tokenizer. Anthropic reports that identical input text produces approximately 30% more tokens than on Haiku 4.5, with the actual increase depending on the content. Consequently, old token counts are not a reliable way to estimate the new model's prompt length or cost.[9]
The token-counting endpoint accepts a model ID and structured message input, including supported tool definitions, system prompts, images, and PDFs. Counts are estimates and can differ slightly from a generation's actual usage. Images and PDFs must use base64 rather than URL or file sources for that endpoint; most server tools and the MCP connector are not accepted. For those requests, the Messages API's response reports actual usage.[9]
Migration from Haiku 4.5
Haiku 5.5 changes API behavior as well as the model name. Its migration guide identifies five breaking changes:[7]
| Haiku 4.5 integration pattern | Haiku 5.5 requirement |
|---|---|
Manual thinking with budget_tokens | Use adaptive thinking and effort |
| Non-default sampling parameters | Omit temperature, top_p, and top_k |
| Final assistant message used as a prefill | End with a user turn; use instructions, tools, or structured outputs for formatting |
computer_20250124 on the Claude API or Google Cloud | Use computer_toolset_20260801 |
| Editing earlier context while replaying thinking | Keep the conversation prefix unchanged |
For example, a Messages API request can explicitly select adaptive thinking and an effort level:[7]
{
"messages": [{"role": "user", "content": "Summarize this support ticket: the invoice contains a duplicate charge."}],
"output_config": {"effort": "medium"},
"thinking": {"type": "adaptive"},
"max_tokens": 4096,
"model": "claude-haiku-5-5"
}
Managed Agents requires only a model-name update. Haiku 5.5 does not support Priority Tier.[7]
Computer and browser tools
On the Claude API and Google Cloud, Haiku 5.5 supports the computer toolset and the browser_toolset_20260801 browser toolset. Haiku 4.5 does not support the latter.[10]
Replaying signed thinking
Haiku 5.5 thinking blocks are bound to the account that produced them, or an account linked to it. Replaying them through another account succeeds, but the API drops those blocks and the model answers without their reasoning. With the thinking-binding-controls-2026-08-01 beta header on the Claude API or Google Cloud, input_transformations identifies these drops as organization_binding_mismatch.[11]
A block also depends on the unchanged system prompt, tools, and messages preceding it. Prefix checking is enforced by default for accounts created on or after August 31, 2026, 00:00 UTC. Older accounts enable enforcement by setting thinking.block_binding.prefix_mismatch_behavior. Keeping a session append-only avoids relying on the account's age. Server-side compaction changes where the checked prefix begins.[11]
Prompting and application behavior
Anthropic's Haiku-specific prompting guide documents several behaviors relevant to deployment. At low effort with long agent prompts, the model can stop early or omit a needed search. At low or medium effort, it can report a code change complete without exercising it. The guide recommends explicit completion and verification instructions, with higher effort as another option at greater token cost.[12]
For search-backed answers, Anthropic recommends supplying the current date and targeted instructions to search for facts that may have changed. For structured output combined with application tools, disabling thinking can cause a needed tool call to be skipped; adaptive thinking or a forced tool choice can address this behavior. Mid-task user messages should arrive as user content, not be embedded in untrusted tool results. At xhigh in multi-turn chats, the model can sometimes produce an answer in thinking without visible text, so clients should check for empty replies.[12]
Safeguards and refusals
Classifier refusals are successful HTTP 200 responses with stop_reason: "refusal", rather than transport errors. stop_details describes the category and explanation; its explanatory text is not a stable machine-readable identifier. A refusal can occur before output or after partial streamed output. Partial output from a refused request is incomplete and should be discarded.[13]
Haiku 5.5 has no server-side fallback. Setting fallbacks: "default" does not turn a declined request into a successful one, and supplying a fallback-model list returns a 400 error. Clients must handle the refusal themselves. Some refusals can be billed, including pre-output bio, frontier_llm, and reasoning_extraction refusals; benign requests can also trigger classifiers.[13]
Anthropic's pre-deployment assessment found improved prompt-injection resistance over Haiku 4.5, but also more over-refusal than any comparison model in its automated behavioral audit. These are evaluation findings, not a guarantee that an agent will resist every attack or accept every legitimate request.[14]
Reported evaluations
The system card reports the following Haiku 5.5 results. Its summary uses adaptive thinking at max effort and averages five trials unless the individual evaluation specifies otherwise.[14]
| Evaluation | Reported result |
|---|---|
| SWE-Bench Pro | 64.8% |
| FrontierCode 1.1 Main | 46.4% at max |
| Humanity's Last Exam, no tools | 45.9% |
| Humanity's Last Exam, with tools | 57.4% |
| OSWorld 2.1 offline subset | 72.4% partial credit; 37.1% strict pass |
| Terminal-Bench 4.0 | 39.2% |
| GDPval-AA v2.1 | 1620 Elo |
| AA-Briefcase v1.1 | 1578 Elo |
OSWorld used 82 offline tasks, not all 108 tasks. Terminal-Bench used 66 tasks and ten trials each, with safeguards enabled, no fallback model, and no internet egress. Twelve blocked trials failed. These settings differ from ordinary default-medium use.[14]
The original SWE-Bench Pro paper defines repository-level software tasks with human-reviewed issue requirements and tests, including tests that check both the requested fix and preservation of existing behavior. Its public, held-out, and commercial sets have different access policies. A score on it therefore concerns a particular evaluated software-agent setup, rather than an all-purpose coding accuracy percentage.[15]
The OSWorld 2.0 paper distinguishes progress through checkpoints from completing the entire workflow. It evaluates long tasks across applications and stateful services, including missing information and changes that arrive during execution. This distinction explains why partial credit and strict pass rate should remain separate when interpreting Haiku's reported computer-use results.[16]
References
- ^1 ^2Anthropic. "Claude Haiku 5.5". October 7, 2026. Accessed October 8, 2026.
- ^1 ^2 ^3Anthropic, Claude Platform Docs. "Claude Haiku 5.5". Accessed October 8, 2026.
- ^1 ^2Anthropic, Claude Platform Docs. "Prompt caching". Accessed October 8, 2026.
- ^1 ^2 ^3Anthropic, Claude Platform Docs. "Batch processing". Accessed October 8, 2026.
- ^1 ^2Anthropic, Claude Platform Docs. "Steering thinking". Accessed October 8, 2026.
- ^1 ^2 ^3Anthropic, Claude Platform Docs. "Effort". Accessed October 8, 2026.
- ^1 ^2 ^3 ^4Anthropic, Claude Platform Docs. "Claude Haiku 5.5 migration guide". Accessed October 8, 2026.
- ^1 ^2 ^3Anthropic, Claude Platform Docs. "Context windows". Accessed October 8, 2026.
- ^1 ^2Anthropic, Claude Platform Docs. "Token counting". Accessed October 8, 2026.
- ^1 ^2 ^3Anthropic, Claude Platform Docs. "What's new in Claude Haiku 5.5". Accessed October 8, 2026.
- ^1 ^2Anthropic, Claude Platform Docs. "Preserved thinking". Accessed October 8, 2026.
- ^1 ^2Anthropic, Claude Platform Docs. "Prompting Claude Haiku 5.5". Accessed October 8, 2026.
- ^1 ^2Anthropic, Claude Platform Docs. "Refusals and fallback". Accessed October 8, 2026.
- ^1 ^2 ^3Anthropic. "Claude Haiku 5.5 System Card". October 7, 2026, sections 5-6 and 8. Accessed October 8, 2026.
- ^Deng, Xiang, et al. "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?". arXiv:2509.16941, 2025, sections 3-4. Accessed October 8, 2026.
- ^Yuan, Mengqi, et al. "OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks". arXiv:2606.29537, 2026, section 2. Accessed October 8, 2026.
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 2,223 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent full-article review against 16 primary and research references, October 8, 2026. Checked subject identity, specifications, availability, limitations and citation support.
Cite this page: AI Wiki. "Claude Haiku 5.5." aiwiki.ai, updated 7 Oct 2026, fact-checked 7 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/claude_haiku_5_5