Citation and evidence

Claude Haiku 5.5

11 min full readUpdated 16 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI ModelsAI SafetyAnthropicLarge Language Models

Cite this article

Claude Haiku 5.5 is a large language model released by Anthropic on October 7, 2026. It belongs to the small, fast tier of the Claude family. Anthropic positions it for high-volume work such as summarization, classification, database queries, customer support, and narrowly scoped coding subagents.[1]

Its API provides a one-million-token context window, a standard output limit of 128,000 tokens, and adaptive thinking. The default effort is medium.[10][6]

Model specifications and availability

PropertyClaude Haiku 5.5
Release dateOctober 7, 2026
Claude API model IDclaude-haiku-5-5
Input and outputText and images to text
Context window1,000,000 tokens
Standard maximum output128,000 tokens
Reliable knowledge cutoffJune 2026
Training data cutoffJune 2026
Default effortmedium
Retirement commitmentNot before October 7, 2027

Expanded article table

These are the published API specifications.[2]

PlatformPublished model ID
Claude APIclaude-haiku-5-5
Amazon Bedrockanthropic.claude-haiku-5-5
Claude Platform on AWSclaude-haiku-5-5
Google Cloudclaude-haiku-5-5
Microsoft Foundryclaude-haiku-5-5

Expanded article table

The platform IDs above are listed in the model overview.[2]

The Claude API ID is fixed, without a date suffix or separate alias. Haiku 4.5 integrations need their platform's replacement ID.[7]

Pricing and prompt length

Haiku 5.5 has two prompt-length pricing tiers. The following Claude API list prices are in US dollars per million tokens, as listed on October 8, 2026. A prompt longer than 100,000 tokens uses the higher tier, including the higher output rate.[1]

Token categoryPrompt at or below 100,000 tokensPrompt above 100,000 tokens
Uncached input$0.10$0.50
Output$0.50$2.50
Five-minute cache write$0.125$0.625
One-hour cache write$0.20$1.00
Cache read$0.01$0.05
Batch input$0.05$0.25
Batch output$0.25$1.25 [2][4]

Expanded article table

Prompt caching

Prompt caching reuses an unchanged prefix containing tool definitions, a system prompt, and conversation content. A cache hit reduces the price of those input tokens; it does not remove them from the request. The response distinguishes uncached input_tokens, cache_creation_input_tokens, and cache_read_input_tokens.[3]

For Haiku 5.5, the published minimum cacheable prefix is 512 tokens on the Claude API, Claude Platform on AWS, Google Cloud, and Microsoft Foundry. A shorter marked prefix runs without caching rather than producing a caching error. Bedrock has its own caching documentation and usage fields. Five-minute entries refresh when reused; a one-hour duration is available by setting ttl: "1h" in cache_control.[3]

Changing top-level effort between requests invalidates cached conversation messages. Haiku 5.5 supports a beta per-message effort change on the Claude API and Google Cloud that preserves the earlier cache. It requires adaptive thinking and the mid-conversation-output-config-2026-07-01 beta header; it cannot change effort while thinking is disabled.[6]

Batch processing and extended output

The Message Batches API processes requests asynchronously at half the standard input and output rates. It is suited to work that does not require an immediate answer. Each batch is limited to 100,000 requests or 256 MB, whichever comes first, and unfinished requests expire after 24 hours.[4]

For Haiku 5.5, the output-300k-2026-03-24 beta header raises the batch output limit to 300,000 tokens. This is not the synchronous Messages API's standard 128,000-token limit. The extended-output beta is available on the Claude API and Claude Platform on AWS, not Amazon Bedrock, Google Cloud, or Microsoft Foundry. A single large generation can take more than an hour.[4]

Thinking and effort

Adaptive thinking is enabled by default. The model decides how much reasoning to allocate to a request and can reason between tool calls. An effort level is a behavioral control, not a guaranteed token budget. max_tokens is the hard per-request limit on thinking plus the answer, so a small value can leave no room for visible text.[5]

Thinking tokens are billed as output even when their text is omitted from the response. usage.output_tokens_details.thinking_tokens reports the raw reasoning-token count; output_tokens remains the inclusive billed total. A tool-use loop makes multiple requests, each with its own output cap, so one request's max_tokens does not cap the cost of the entire task.[5]

EffortDocumented starting use
lowShort tool tasks, chat, and simple high-volume requests
mediumDefault; most work, including agentic coding
highKnowledge work, longer tasks, and stricter instruction following
xhighWork where application evaluations justify additional reasoning
maxWork where application evaluations justify the highest token expenditure

Expanded article table

Anthropic recommends comparing higher-effort Haiku configurations with Claude Sonnet 5.5 on quality, latency, and cost. Haiku permits thinking: {"type": "disabled"} at low, medium, or high; disabling thinking at xhigh or max returns a 400 error.[6]

The response can begin with a thinking block rather than answer text. Thinking text is omitted by default. To receive summarized reasoning, set thinking.display to "summarized". Applications displaying the answer should select text blocks by type instead of assuming the first block is text.[10]

Context and token accounting

The context window covers the system prompt, tool definitions, messages, tool results, image and document tokens, and the output being generated. Cached prefixes still occupy it. Haiku 5.5 preserves earlier thinking blocks by default, unlike Haiku 4.5, which strips previous thinking blocks from subsequent context.[8]

The one-million-token window is the default and does not need a long-context beta header. Requests can contain up to 600 images or PDF pages, but request-size limits can be reached before the token limit. Input that already exceeds the window returns a 400 error. If generation reaches the window limit, the response can stop with stop_reason: "model_context_window_exceeded".[8]

Haiku 5.5 does not receive the injected remaining-context-budget tags used by Haiku 4.5. Developers can supply an explicit task budget through the task-budgets beta instead.[8]

Haiku 5.5 uses the newer Claude tokenizer. Anthropic reports that identical input text produces approximately 30% more tokens than on Haiku 4.5, with the actual increase depending on the content. Consequently, old token counts are not a reliable way to estimate the new model's prompt length or cost.[9]

The token-counting endpoint accepts a model ID and structured message input, including supported tool definitions, system prompts, images, and PDFs. Counts are estimates and can differ slightly from a generation's actual usage. Images and PDFs must use base64 rather than URL or file sources for that endpoint; most server tools and the MCP connector are not accepted. For those requests, the Messages API's response reports actual usage.[9]

Migration from Haiku 4.5

Haiku 5.5 changes API behavior as well as the model name. Its migration guide identifies five breaking changes:[7]

Haiku 4.5 integration patternHaiku 5.5 requirement
Manual thinking with budget_tokensUse adaptive thinking and effort
Non-default sampling parametersOmit temperature, top_p, and top_k
Final assistant message used as a prefillEnd with a user turn; use instructions, tools, or structured outputs for formatting
computer_20250124 on the Claude API or Google CloudUse computer_toolset_20260801
Editing earlier context while replaying thinkingKeep the conversation prefix unchanged

Expanded article table

For example, a Messages API request can explicitly select adaptive thinking and an effort level:[7]

{
  "messages": [{"role": "user", "content": "Summarize this support ticket: the invoice contains a duplicate charge."}],
  "output_config": {"effort": "medium"},
  "thinking": {"type": "adaptive"},
  "max_tokens": 4096,
  "model": "claude-haiku-5-5"
}

Managed Agents requires only a model-name update. Haiku 5.5 does not support Priority Tier.[7]

Computer and browser tools

On the Claude API and Google Cloud, Haiku 5.5 supports the computer toolset and the browser_toolset_20260801 browser toolset. Haiku 4.5 does not support the latter.[10]

Replaying signed thinking

Haiku 5.5 thinking blocks are bound to the account that produced them, or an account linked to it. Replaying them through another account succeeds, but the API drops those blocks and the model answers without their reasoning. With the thinking-binding-controls-2026-08-01 beta header on the Claude API or Google Cloud, input_transformations identifies these drops as organization_binding_mismatch.[11]

A block also depends on the unchanged system prompt, tools, and messages preceding it. Prefix checking is enforced by default for accounts created on or after August 31, 2026, 00:00 UTC. Older accounts enable enforcement by setting thinking.block_binding.prefix_mismatch_behavior. Keeping a session append-only avoids relying on the account's age. Server-side compaction changes where the checked prefix begins.[11]

Prompting and application behavior

Anthropic's Haiku-specific prompting guide documents several behaviors relevant to deployment. At low effort with long agent prompts, the model can stop early or omit a needed search. At low or medium effort, it can report a code change complete without exercising it. The guide recommends explicit completion and verification instructions, with higher effort as another option at greater token cost.[12]

For search-backed answers, Anthropic recommends supplying the current date and targeted instructions to search for facts that may have changed. For structured output combined with application tools, disabling thinking can cause a needed tool call to be skipped; adaptive thinking or a forced tool choice can address this behavior. Mid-task user messages should arrive as user content, not be embedded in untrusted tool results. At xhigh in multi-turn chats, the model can sometimes produce an answer in thinking without visible text, so clients should check for empty replies.[12]

Safeguards and refusals

Classifier refusals are successful HTTP 200 responses with stop_reason: "refusal", rather than transport errors. stop_details describes the category and explanation; its explanatory text is not a stable machine-readable identifier. A refusal can occur before output or after partial streamed output. Partial output from a refused request is incomplete and should be discarded.[13]

Haiku 5.5 has no server-side fallback. Setting fallbacks: "default" does not turn a declined request into a successful one, and supplying a fallback-model list returns a 400 error. Clients must handle the refusal themselves. Some refusals can be billed, including pre-output bio, frontier_llm, and reasoning_extraction refusals; benign requests can also trigger classifiers.[13]

Anthropic's pre-deployment assessment found improved prompt-injection resistance over Haiku 4.5, but also more over-refusal than any comparison model in its automated behavioral audit. These are evaluation findings, not a guarantee that an agent will resist every attack or accept every legitimate request.[14]

Reported evaluations

The system card reports the following Haiku 5.5 results. Its summary uses adaptive thinking at max effort and averages five trials unless the individual evaluation specifies otherwise.[14]

EvaluationReported result
SWE-Bench Pro64.8%
FrontierCode 1.1 Main46.4% at max
Humanity's Last Exam, no tools45.9%
Humanity's Last Exam, with tools57.4%
OSWorld 2.1 offline subset72.4% partial credit; 37.1% strict pass
Terminal-Bench 4.039.2%
GDPval-AA v2.11620 Elo
AA-Briefcase v1.11578 Elo

Expanded article table

OSWorld used 82 offline tasks, not all 108 tasks. Terminal-Bench used 66 tasks and ten trials each, with safeguards enabled, no fallback model, and no internet egress. Twelve blocked trials failed. These settings differ from ordinary default-medium use.[14]

The original SWE-Bench Pro paper defines repository-level software tasks with human-reviewed issue requirements and tests, including tests that check both the requested fix and preservation of existing behavior. Its public, held-out, and commercial sets have different access policies. A score on it therefore concerns a particular evaluated software-agent setup, rather than an all-purpose coding accuracy percentage.[15]

The OSWorld 2.0 paper distinguishes progress through checkpoints from completing the entire workflow. It evaluates long tasks across applications and stateful services, including missing information and changes that arrive during execution. This distinction explains why partial credit and strict pass rate should remain separate when interpreting Haiku's reported computer-use results.[16]

References

  1. ^1 ^2Anthropic. "Claude Haiku 5.5". October 7, 2026. Accessed October 8, 2026.
  2. ^1 ^2 ^3Anthropic, Claude Platform Docs. "Claude Haiku 5.5". Accessed October 8, 2026.
  3. ^1 ^2Anthropic, Claude Platform Docs. "Prompt caching". Accessed October 8, 2026.
  4. ^1 ^2 ^3Anthropic, Claude Platform Docs. "Batch processing". Accessed October 8, 2026.
  5. ^1 ^2Anthropic, Claude Platform Docs. "Steering thinking". Accessed October 8, 2026.
  6. ^1 ^2 ^3Anthropic, Claude Platform Docs. "Effort". Accessed October 8, 2026.
  7. ^1 ^2 ^3 ^4Anthropic, Claude Platform Docs. "Claude Haiku 5.5 migration guide". Accessed October 8, 2026.
  8. ^1 ^2 ^3Anthropic, Claude Platform Docs. "Context windows". Accessed October 8, 2026.
  9. ^1 ^2Anthropic, Claude Platform Docs. "Token counting". Accessed October 8, 2026.
  10. ^1 ^2 ^3Anthropic, Claude Platform Docs. "What's new in Claude Haiku 5.5". Accessed October 8, 2026.
  11. ^1 ^2Anthropic, Claude Platform Docs. "Preserved thinking". Accessed October 8, 2026.
  12. ^1 ^2Anthropic, Claude Platform Docs. "Prompting Claude Haiku 5.5". Accessed October 8, 2026.
  13. ^1 ^2Anthropic, Claude Platform Docs. "Refusals and fallback". Accessed October 8, 2026.
  14. ^1 ^2 ^3Anthropic. "Claude Haiku 5.5 System Card". October 7, 2026, sections 5-6 and 8. Accessed October 8, 2026.
  15. ^Deng, Xiang, et al. "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?". arXiv:2509.16941, 2025, sections 3-4. Accessed October 8, 2026.
  16. ^Yuan, Mengqi, et al. "OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks". arXiv:2606.29537, 2026, section 2. Accessed October 8, 2026.

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

v1 · 2,223 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independent full-article review against 16 primary and research references, October 8, 2026. Checked subject identity, specifications, availability, limitations and citation support.

Cite this page: AI Wiki. "Claude Haiku 5.5." aiwiki.ai, updated 7 Oct 2026, fact-checked 7 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/claude_haiku_5_5

Suggest edit

What links here