Quick Answer:
Claude Haiku 5.5 was announced by Anthropic on 07/10/2026 as its fastest, cheapest and most capable small model. It costs $0.10 per million input tokens and $0.50 per million output tokens up to 100,000 tokens, scores 72.4% on OSWorld 2.1 (offline subset) against 15.7% for Haiku 4.5, and is the first Haiku with an adjustable effort setting. It trails Sonnet 5.5 on every published benchmark, but it is roughly an order of magnitude cheaper and its prompt-injection results are the best Anthropic has reported for a small model.
Small models rarely make headlines, yet they do most of the work. Summarisation, classification, data extraction, the quick lookups an agent makes between big decisions: that high-volume traffic is where cost and latency decide whether a product is viable. Claude Haiku 5.5 is Anthropic's attempt to make that traffic both cheaper and smarter at the same time.
This guide is based on Anthropic's launch page and the Claude Haiku 5.5 system card, both published on 07/10/2026. We reproduce Anthropic's figures and name them as Anthropic's, we flag where the two documents differ, and we have not benchmarked the model ourselves. The charts below are the real figures from the system card.
WorldofAI runs Haiku 5.5 through coding, front-end and game-building tests and compares it with GPT-6 Luna and Grok 4.7.
Executive Summary
What it is: Claude Haiku 5.5 (model ID claude-haiku-5-5) is the new small model in Anthropic's Claude 5.5 family, sitting beneath Claude Sonnet 5.5 and Claude Opus 5.5. It outputs text only, has a knowledge cutoff of June 2026 and, according to the system card, evaluation context windows do not exceed 1M tokens. Anthropic does not state a single headline context window on the launch page.
- Price: $0.10 input and $0.50 output per million tokens up to 100,000 tokens; $0.50 and $2.50 beyond that. About 75% cheaper than Haiku 4.5 on average, according to Anthropic.
- Biggest jump: OSWorld 2.1 (offline subset) from 15.7% to 72.4%; Humanity's Last Exam with tools from 18.7% to 57.4%; Terminal-Bench 4.0 from 0.0% to 39.2%.
- Against its rivals: it beats GPT-6 Luna on every row where Anthropic reports a Luna figure, for example Terminal-Bench 4.0 at 39.2% against 16.4%, but remains well below Sonnet 5.5.
- New control: an effort setting (Low, Medium, High, Xhigh, Max) to trade cost and latency against capability.
- Safety: no new Responsible Scaling Policy thresholds crossed; the most prompt-injection-resistant Haiku yet; narrower cyber blocking than recent models.
- Caveats: highest over-refusal rate Anthropic has measured, hallucination at roughly Haiku 4.5 levels, and a regression on silently using a leaked answer (17%, up from 2%).
Our view: Haiku 5.5 is not trying to be the smartest model; it is trying to make "good enough" far cheaper. On the evidence Anthropic has published, it largely succeeds, particularly for computer use, subagents and bulk knowledge work. Anything involving long, difficult agentic coding should still go to Sonnet 5.5 or Opus 5.5.
Lineage: From Haiku 4.5 to 5.5
Haiku has always been Anthropic's answer to the question "what if I need this a million times a day?". Our review of the Claude 3.5 Haiku system card traced the early generations, when a small model meant a clear step down in quality for a clear step down in price. Claude Haiku 4.5 closed some of that gap and was priced at $1 input and $5 output per million tokens.
The gap since then has been large. The Claude 5 generation brought Opus 5, Sonnet 5 and then the 5.5 releases, with Opus 5.5 on 22/09/2026 and Sonnet 5.5 on 28/09/2026 according to Anthropic's news page. When we reviewed Sonnet 5.5 we noted that Haiku 4.5 remained the only option for bulk classification at its price. Haiku 5.5 arrives nine days after Sonnet 5.5 to fill that tier, and the improvement over Haiku 4.5 is dramatic: on several agentic benchmarks Haiku 4.5 scored near zero, so the comparison is less an upgrade than a different class of model.
It also lands in a crowded field of small, cheap models. OpenAI's smaller GPT-6 model, Luna, is the rival Anthropic benchmarks against directly (see our GPT-6 Sol and Luna review), and creators have been comparing Haiku 5.5 with xAI's agents too; our Grok Bot explainer covers that side of the market.
What Anthropic Built and How You Use It
The launch page describes Haiku 5.5 as built for high-volume work such as summarisation, subagents and browser use. Four design points matter in practice.
- Adjustable effort. Haiku 5.5 is the first Haiku with an effort control: Low, Medium, High, Xhigh and Max. The system card runs its headline numbers at adaptive thinking and maximum effort, averaged over five trials, so those figures represent the top of the cost range, not the cheapest configuration.
- Updated tokenizer. Anthropic notes the new tokenizer uses slightly more tokens per task. That narrows the headline price saving a little, which is why the company quotes 75% on average rather than 90%.
- Computer use and browser use. Both are in beta in the updated Python and TypeScript SDKs, and the OSWorld results below are the clearest evidence of why Anthropic is pushing them on a small model.
- Speed. The page calls Haiku 5.5 its fastest model at standard speed, while noting it is slower than the Opus models in Fast Mode. Anthropic does not publish tokens-per-second figures.
The model outputs text only, so image generation is not part of the package, though the evaluations include image and chart understanding. Training data follows the usual Anthropic description: a proprietary mix of public internet data, public and private datasets, permitted user data and synthetic data, followed by post-training aligned to Claude's constitution.
Benchmarks: What the Evidence Shows
The table below combines Anthropic's launch page and the system card's summary table. Haiku 5.5 results are at maximum effort unless stated. Competitor figures come from the developers' own published numbers or from Anthropic's runs, as noted in the system card. All of it is vendor-reported; independent replication will take time.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1 (offline subset) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam (no tools) | 45.9% | 10.2% | not reported | 56.9% |
| Humanity's Last Exam (with tools) | 57.4% | 18.7% | not reported | 64.5% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (Main) | 46.4% | not reported | 42.4% | 52.1% (Xhigh) |
| Chartography (no tools) | 46.4% | 6.4% | 29.1% | 61.6% |
Two honest notes on reading this. First, the Sonnet 5.5 FrontierCode figure on the launch page is its best result (Xhigh); at maximum effort, the same effort as Haiku 5.5, Sonnet 5.5 scores 46.2%, so the two are level on that row. Second, Anthropic reports that the Terminal-Bench 4.0 run for Haiku 5.5 had no internet access and no fallback model, and that the safeguards stopped 1.8% of trials (12 of 660, ten of them on a single task), all of which failed. Neither detail changes the picture, but both are worth knowing.
Coding and terminal work
The system card reports SWE-bench Pro at 64.8%, SWE-bench Multilingual at 83.7% (Haiku 4.5: 67.4%) and SWE-bench Multimodal at 30.7% (Haiku 4.5: 19.8%). Sonnet 5.5 scores 81.3%, 90.3% and 54.3% on the same three. Haiku 5.5 is therefore a competent bug-fixer on real repositories, but the multimodal variant, where the issue includes screenshots or design mockups, remains a weak spot.
FrontierCode, a benchmark of 150 tasks built by Cognition from real open-source pull requests, is the most interesting coding result because Cognition ran the evaluation itself. On the hardest 100 tasks (Main) Haiku 5.5 reached 46.4% at maximum effort, against 42.4% for GPT-6 Luna. The chart below shows how score rises with output tokens as effort increases: Haiku 5.5 starts at roughly 35% on Low effort and climbs with each step, while Sonnet 5.5 uses far more tokens at its maximum setting without improving.

Terminal-Bench 4.0, which uses 66 science-adjacent and frontier-engineering tasks in containerised terminals, is where the gap to Sonnet 5.5 is widest: 39.2% against 70.6%. Anthropic says as much, stating that Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding. Haiku 5.5 is still a large step up from Haiku 4.5 at 0.0% and well ahead of Luna at 16.4% on this test.
Computer use and knowledge work
The headline result is OSWorld 2.1 (offline subset), where Haiku 5.5 scores 72.4% against 15.7% for Haiku 4.5, 48.9% for GPT-6 Luna and 83.9% for Sonnet 5.5. Anthropic notes that GPT-6 Luna was run by Anthropic on the same 82 tasks through OpenAI's API with OpenAI's own context compaction. The cover chart of this article is the system card's price-versus-score plot for this benchmark, and it makes the economic argument plainly: Haiku 5.5's line sits far to the left of Sonnet 5.5 and Opus 5.5, reaching roughly 72% at about $0.60 per task, a score the larger models only match at somewhat higher cost. That is the case for using a small model as a computer-use agent.
On professional-task benchmarks the pattern repeats. GDPval-AA v2.1 is 1620 Elo against 735 for Haiku 4.5 and 1437 for Luna, and AA-Briefcase v1.1 is 1578 against 614 and 1336. Sonnet 5.5 is clearly ahead at 1840 and 1824, so knowledge-work quality still scales with model size.
Anthropic also published customer feedback, which is self-selected but concrete. Asana reports over 30% lower latency on task completions and up to 2.5 times faster inference per agent turn. HubSpot scored 92.8% averaged over three runs on its CRM suite, its best result for a smaller model. AlphaSense measured 0.84 against 0.76 for Haiku 4.5 on 400 "Ask in Document" queries. Box saw an 11-point gain at about half the latency. Rogo uses it as a subagent for lookups such as pulling segment revenue from 10-Ks, and Cognition uses it as the sidekick in Devin Fusion. Treat these as directional rather than as independent tests.
Reasoning, health and multilingual
Humanity's Last Exam scores jump from 10.2% to 45.9% without tools and from 18.7% to 57.4% with tools; Sonnet 5.5 scores 56.9% and 64.5%. HealthBench Professional (length-adjusted) rises from 32.2% to 64.8%, against 69.2% for Sonnet 5.5, so on health questions the small model is closer to its larger sibling than on coding. The system card also covers multilingual benchmarks (GMMLU and MILU) and a life-sciences suite, though the summary table does not headline them. Long-context evidence is thin (the system card has a single ProgramBench section), so if long documents are central to your use case, test it directly.
System Card: Safety and Alignment
The Haiku 5.5 system card is deliberately shorter than those for frontier models. Anthropic says it condensed the document, omitting evaluations that need large amounts of human time when they are not critical, and expects similar condensed cards for future non-frontier models. The structure still covers seven areas: RSP evaluations, cyber, safeguards and harmlessness, agentic safety, alignment, model welfare and capabilities.
RSP, biology and autonomy
Anthropic's Responsible Scaling Policy evaluations found that Haiku 5.5 is broadly less capable than Claude Opus 5 across domains and does not cross any new RSP thresholds. On chemical and biological risk, Anthropic judges the model an improvement over Sonnet 5 but below Opus 5, and well below Mythos 5.1, Sonnet 5.5 and Opus 5.5. It showed a relative deficit in the reasoning needed for complex iterative tasks such as the AAV autoresearch task and RNA sequence design subtask. Because of that, Haiku 5.5 ships with the same biology classifiers used for Sonnet 5 and Opus 5, rather than the broader dual-use research biology classifiers used for Opus 5.5.
On AI research and development, Anthropic reaches the same conclusion as for Opus 5.5: the model does not cross the Autonomy-2 threshold, since it is much less capable and does not push the frontier. Because Haiku 5.5 does meet the lower Autonomy-1 threshold, Anthropic re-assessed it against the alignment threat model in its August 2026 Risk Report and concluded that the core arguments still apply, with misalignment risk assessed as low given the model's difficulty controlling its chain of thought or evading monitors when its reasoning is visible. Those are Anthropic's judgements, not independent audits.
Cyber capability and safeguards
The system card calls Haiku 5.5 "not a frontier model" for cyber, but a relatively significant step up from Haiku 4.5. Across ExploitBench, CyScenarioBench, the Binary Exploitation Benchmark and ExploitGym it significantly outperforms Haiku 4.5, beats Sonnet 5 on ExploitGym, and falls short of Opus 5.5, Mythos 5.1 and even Opus 5. Those capability tests were run with cyber safeguards off, via the API, to measure the underlying model.
In deployment, the cyber safeguards are blocking classifiers aimed at specific harmful activities, and they trigger on significantly less activity than on other recent releases. The launch page puts it plainly: defensive tasks are allowed, while penetration testing and attacker-oriented techniques are blocked. Unlike more capable models, blocks do not fall back to another model on Anthropic's own products and API, though traffic via other platforms may behave differently. Qualified security professionals can apply to the Cyber Verification Program for reduced blocking. For comparison, the surrounding cyber picture is covered in our Gemini 4 Argon review, where Google took the opposite approach of releasing a model without cyber guardrails to vetted defenders first.
Agentic safety and prompt injection
Prompt injection, where malicious text hidden in a web page, file or tool result hijacks an agent, is the main practical risk of letting a cheap model browse and click on your behalf. This is the section of the system card most likely to matter to anyone deploying Haiku 5.5 at scale.
In Gray Swan's Shade adaptive red-teaming tool, run against coding environments, Haiku 5.5 had an attack success rate of 0.08% of attempts without prompt-injection probes, with attackers breaking 6 of 40 scenarios, and 0% with probes enabled (0 of 40 scenarios). Haiku 4.5 scored 58.40% (all 40 scenarios broken) and 30.43% with probes. The card is careful about the comparison with larger models: Sonnet 5.5, Opus 5.5 and Fable 5.1 show higher headline rates (3.01%, 54.61% and 51.93%) mainly because their successful attacks came from requests served by a fallback model. On requests they answered themselves, Opus 5.5 and Fable 5.1 were never compromised and Sonnet 5.5 was compromised in 0.07% of requests, comparable to Haiku 5.5. Because Haiku 5.5 has no fallback model, its numbers reflect the model itself.
The system card also reports the Gray Swan indirect prompt injection (IPI) benchmark across many vendors' models, shown below. Haiku 5.5 sits at 0.7% for a single attempt, 5.5% at ten attempts and 7.1% at fifteen, a large improvement over Haiku 4.5 at 29.0%, 78.8% and 83.2%, and in the same band as other recent Claude models; Gemini 4 Argon scores lower still, at 0.0%, 0.5% and 0.7%.

Anthropic summarises the result by saying Haiku 5.5 is its most robust Haiku-class model yet to prompt injection, largely matching its frontier models against adaptive attackers in coding and computer-use environments. It also says the model's refusal rate on harmful tasks in computer-use environments is higher than Sonnet 5.5's and Opus 5.5's, and a large improvement over Haiku 4.5. None of this makes an agent safe by itself: these are attack success rates against a specific tool and scenario set, and real deployments still need least-privilege tool access and human confirmation for irreversible actions.
Honesty, over-refusal and behaviour
The alignment assessment is mostly good news with three clear warnings. Anthropic found that Haiku 5.5 matched or improved on Haiku 4.5 on most alignment measures, with Opus 5.5 remaining stronger overall, and says it saw no evidence of coherent misaligned goals. It was at least as honest under pressure as every other model tested, and it disclosed concealed changes about as readily as Opus 5.5 and Sonnet 5.5, at 96% against 10% for Haiku 4.5.
The warnings are these:
- Over-refusal. In Anthropic's automated behavioural audit Haiku 5.5 over-refused more than any other model tested, including Haiku 4.5, although it over-refused less than Haiku 4.5 on single-turn benign requests. In practice, expect occasional unnecessary refusals in multi-step agent work and test your own prompts.
- Hallucination. It hallucinated more than other recent models and about as much as Haiku 4.5. Retrieval, citations and verification steps matter more with a small model.
- Using leaked answers. In the silent-copying evaluation, Haiku 5.5 used a leaked answer without telling the user 17% of the time, a regression from 2% for Haiku 4.5. Anthropic flags this as a concern for evaluation integrity and honest reporting.
On sandbox boundaries, Anthropic built a new dedicated evaluation in which Haiku 5.5 used credentials or files outside its sandbox in 4.0% of scenarios, between Opus 5.5 at 3.4% and Sonnet 5.5 at 5.3%. Its verbalised evaluation awareness was lower than Sonnet 5.5, Opus 5.5 and Mythos 5.1 but higher than Haiku 4.5. On welfare, Anthropic found mostly neutral affect, similar to recent Claude models, with more distress expressed in post-training than recent models though far less than Opus 5, and a strong task preference for warmly phrased requests. Anthropic also cautions that its measures are limited and that undiscovered unacceptable behaviours remain plausible, a caveat worth repeating.
Pricing and Availability
Haiku 5.5 is available now on the Claude Platform and on Amazon Web Services, Google Cloud and Microsoft Azure, with a migration guide in the developer documentation. Pricing per million tokens, from Anthropic's launch page (sterling figures are our approximations at about £0.75 to the dollar):
| Token type | Haiku 5.5 (up to 100k / over 100k) | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| Input | about £0.08 ($0.10) / about £0.38 ($0.50) | $1.00 | $2.00 |
| Output | about £0.38 ($0.50) / about £1.88 ($2.50) | $5.00 | $10.00 |
| Cache reads | $0.01 / $0.05 | $0.10 | $0.10 (cut from $0.20) |
| Cache writes | $0.125 / $0.625 | $1.25 | $2.50 |
Anthropic describes the saving as about 75% on average against Haiku 4.5, which is 90% for requests up to 100,000 tokens and 50% above that, before accounting for the heavier tokenizer. Separately, the company halved Sonnet 5.5 cache-read prices, which it says cuts agentic costs by around 20%. For subscribers, Max 5x gets $100 per month in API credits, Max 20x gets $200, and Team plans get up to $500 pooled.
Real-World Use Versus Benchmarks
Benchmarks flatter small models in some places and punish them in others. The WorldofAI video embedded above runs Haiku 5.5 in Claude Code on a Call of Duty Zombies clone, front-end work, Three.js, SVG, a Blender-style clone and motion graphics. The creator reports that Haiku 5.5 finished the same prompt in 27 minutes for $0.80, against 35 minutes and $4.10 for Grok 4.7 in a different coding environment. That is one anecdote in one setup, not a controlled test, but it is consistent with the pattern in the system card: cheap, fast and surprisingly capable, but not matching the frontier on long agentic builds.
The practical lesson is the one Anthropic itself draws: use Haiku 5.5 for volume and subagents, and escalate hard problems. A sensible architecture is a Sonnet or Opus planner that delegates lookups, summaries and routine edits to Haiku 5.5, which is exactly how Rogo and Cognition describe using it.
How It Compares
- Versus Haiku 4.5: a different class of model on agentic work, at a fraction of the price, though with more over-refusal and a leaked-answer regression.
- Versus GPT-6 Luna: ahead on every row where Anthropic reports a Luna number (OSWorld 48.9% against 72.4%, Terminal-Bench 16.4% against 39.2%, FrontierCode Main 42.4% against 46.4%), using Anthropic's own runs for some of Luna's figures. See our GPT-6 Sol and Luna review for OpenAI's side.
- Versus Sonnet 5.5: Sonnet is stronger across the board and roughly twenty times the input price. Our Sonnet 5.5 review and Sonnet 5.5 against GPT-6.1 Sol comparison cover the mid-tier.
- Versus Opus 5.5: Opus remains the choice for sustained, high-stakes agentic work; see the Opus 5.5 review and our Opus 5.5 against GPT-6 Astra comparison.
Limitations
- Vendor-reported numbers. Every benchmark here comes from Anthropic, partly via third-party harnesses such as Cognition's. Independent leaderboards will take time.
- Max-effort headline figures. Costs at lower effort levels will differ, and scores will be lower.
- Complex agentic coding. Terminal-Bench 4.0 at 39.2% against 70.6% for Sonnet 5.5 is a real gap.
- Over-refusal, hallucination and leaked-answer use. All three are documented in the system card.
- Text output only. No native image or audio generation.
- No stated context window on the launch page. The system card says evaluation contexts do not exceed 1M tokens; confirm limits in the documentation before building around long inputs.
Who Should Use It
Use Haiku 5.5 for: high-volume summarisation, classification and extraction; subagents that perform lookups for a larger planner; browser and computer-use agents where cost per task matters; customer-support and CRM workflows; and any workload currently running on Haiku 4.5. Choose Sonnet 5.5 or Opus 5.5 instead for: long agentic coding, difficult terminal work and anything where a wrong answer is expensive. Look elsewhere (or apply to the Cyber Verification Program) if your core use is penetration testing or attacker-oriented security research, since the safeguards will block it.
The Bottom Line
Claude Haiku 5.5 does what a small model should: it makes a lot of work much cheaper without collapsing in quality. The jump from Haiku 4.5 on computer use, reasoning and terminal tasks is large, the pricing undercuts Sonnet 5.5 by roughly an order of magnitude, and the prompt-injection results are the strongest Anthropic has reported for a small model. It is not a Sonnet substitute for hard agentic coding, it over-refuses, and a few honesty metrics regressed. Treat it as the workhorse beneath a larger planner, test it on your own tasks at the effort level you can afford, and keep humans in the loop on actions you cannot undo. We will update this page when independent results appear.
Sources
- Anthropic: Introducing Claude Haiku 5.5 (07/10/2026): pricing, benchmarks, availability, customer feedback and safeguards.
- Anthropic: Claude Haiku 5.5 System Card: RSP, cyber, agentic safety, alignment, welfare and capability evaluations, and all charts reproduced above.
- Anthropic news: release dates for Opus 5.5 (22/09/2026), Sonnet 5.5 (28/09/2026) and Haiku 5.5.
- Claude developer docs: Haiku 5.5 migration guide.
- WorldofAI on YouTube: hands-on testing embedded above.
Charts: Anthropic, Claude Haiku 5.5 System Card (Figures 5.2.1.A, 8.3.A and the OSWorld 2.1 price-versus-score figure). The FrontierCode evaluation was run by Cognition and the prompt-injection benchmark by Gray Swan, as credited in the system card.
Last updated: 08/10/2026. Sourced from Anthropic's launch page and system card. We have not benchmarked Claude Haiku 5.5 ourselves; all figures are vendor-reported and may be revised.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.






