Claude Haiku 5.5: Release Date, Price, and Benchmarks
Anthropic released Claude Haiku 5.5 on October 7, 2026. The price, how to get access, the benchmarks vs GPT-6 Luna and Sonnet 5.5, and which effort level to use.

Anthropic released Claude Haiku 5.5 today, October 7, 2026. It costs $0.10 per million input tokens and $0.50 per million output tokens, which is 90% less than Haiku 4.5 and the same sticker price as OpenAI's GPT-6 Luna.
Two weeks ago, in my Claude Opus 5.5 review, I wrote that the one thing missing from Anthropic's lineup was a small, cheap Claude. This is it. But the launch page leaves out a few things you should know before you move your workloads over, and most of them are sitting in the chart data and the system card.
If you only have a minute#
When was it released? October 7, 2026. It is live now, not a staged rollout.
How do you get access? It is in the model picker on Claude.ai for Free, Pro, Max, Team and Enterprise users, in Claude Code, and on the API as claude-haiku-5-5. It is also on Amazon Bedrock, Google Cloud and Microsoft Foundry.
What does it cost? $0.10 in and $0.50 out per million tokens for prompts up to 100,000 tokens. Go over 100,000 and the whole request costs five times more: $0.50 in and $2.50 out.
What is special about it? Three things. It is the first Haiku with an effort setting, so you choose between cheap and smart per request. It runs a computer far better than the last Haiku: 72.4% on OSWorld against 15.7%. And it has a 1 million token context window, up from 200,000.
Where does it fall short? Hard coding in a terminal. Sonnet 5.5 on high effort scores higher than Haiku 5.5 on max and costs 45% less per task. Anthropic's own docs tell you to compare the top two effort levels against Sonnet 5.5 before you use them.
Should you switch? For summaries, classification, extraction, support chat, browser tasks and subagents: yes, today. For the main coding agent: no, keep Sonnet 5.5 or Opus 5.5 there.
I have not run my own tests yet. Everything below comes from Anthropic's announcement, the Claude Platform docs, and the 144-page system card. The cost-per-task numbers are ones I pulled out of the chart data embedded in the announcement page, because Anthropic draws them but never prints them.
How to get access to Claude Haiku 5.5#
This is the part most people are searching for, so here it is plainly.
| Where | How | Cost |
|---|---|---|
| Claude.ai (web, iOS, Android) | Pick Haiku 5.5 in the model picker | Included on Free, Pro, Max, Team and Enterprise |
| Claude Code | Run /model and pick Haiku 5.5 | Counts against your plan, or API billing |
| Claude API | Model ID claude-haiku-5-5 | From $0.10 in, $0.50 out per million tokens |
| Amazon Bedrock | Model ID anthropic.claude-haiku-5-5 | Billed by AWS |
| Google Cloud, Microsoft Foundry | Model ID claude-haiku-5-5 | Billed by the cloud provider |
The plan list comes from Anthropic's Haiku page, and the model IDs come from the model overview. GitHub Copilot and Cursor both published pages for it on launch day too.
If you have never switched models in Claude Code, I wrote a short guide on how to change the model in Claude Code. Medium effort is the default there, the same as on the API.
On the API, a first call looks like this:
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-haiku-5-5",
max_tokens=4096,
output_config={"effort": "low"},
messages=[
{"role": "user", "content": "Classify this ticket as billing, bug, or feature request: ..."}
],
)
for block in response.content:
if block.type == "text":
print(block.text)
Read the answer by block type, like the loop above does. Thinking is on by default, so the first block in a response is not always the text.
What Claude Haiku 5.5 costs#
Here is Anthropic's pricing table. I marked the two rows most people will pay.

| Token type | Up to 100K prompt | Over 100K prompt | Haiku 4.5 |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 |
| Output | $0.50 | $2.50 | $5.00 |
| Cache read | $0.01 | $0.05 | $0.10 |
| Cache write (5 min) | $0.125 | $0.625 | $1.25 |
| Batch input / output | $0.05 / $0.25 | $0.25 / $1.25 | $0.50 / $2.50 |
All prices are per million tokens, from the Claude Platform pricing page.
Three details will change your bill.
The tier is per request, not per token. The system card says each request is charged at the rate for that request's prompt length. A 100,000-token prompt costs one cent of input. Add one token and the same request costs five cents. Every other current Claude model charges a flat rate across its whole 1 million token window, so this is new.
Input cost of one Haiku 5.5 request by prompt length. Cross 100,000 tokens and the whole request is billed at five times the rate, not only the tokens past the line.
Haiku 5.5: $0.008 of input per request, $8.00 per 1,000 requests. Under the line, at the launch price.
Same token count on Haiku 4.5: $0.08. On Sonnet 5.5: $0.16.
Uncached input only, at list prices from the Claude Platform pricing page. Output follows the same two tiers. Haiku 5.5 uses a newer tokenizer that counts about 30% more tokens for the same text than Haiku 4.5, so the same document sits further right on this chart for Haiku 5.5. Word count uses Anthropic’s figure of roughly 555K words per million tokens.
This matters most for agents. A long agent session grows its context with every tool call. Once it passes 100,000 tokens, every later request in that session pays the higher tier. If you use Haiku 5.5 as a subagent, give it a narrow job and a fresh context so it stays under the line.
The same text is more tokens now. Haiku 5.5 uses Anthropic's newer tokenizer, which counts about 30% more tokens for the same text than Haiku 4.5 did. So the real saving on a small prompt is closer to 87% than 90%, and on a big prompt closer to 35% than 50%. That is my arithmetic from Anthropic's two numbers. Anthropic's own blended figure, across real traffic, is "around 75% less" to run.
100,000 tokens is smaller than it sounds. At Anthropic's figure of roughly 555,000 words per million tokens, the cheap tier ends at about 55,000 words of prompt. Anthropic says about 90% of requests to the old Haiku were under that line. Check your own logs before you assume yours are.
Dictating the brief is half the job#
A cheap subagent is only as good as the brief you hand it. A small model with a vague two-line task guesses. The same model with a clear paragraph of context gets it right, and at these prices the longer brief costs you nothing.
I speak most of my long prompts instead of typing them. I built Rambleproof for exactly that: hold a shortcut on your Mac, talk through what you want, and clean text lands wherever your cursor is.
Benchmarks: Haiku 5.5 vs Haiku 4.5, GPT-6 Luna and Sonnet 5.5#
This is the table from Anthropic's launch post. The Haiku 5.5 column highlight is theirs. The two gold boxes are mine: the row where it leads the most, and the row where it trails the most.

| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (office work, Elo) | 1620 | 735 | 1437 | 1840 |
| OSWorld 2.1 (computer use) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam (no tools) | 45.9% | 10.2% | not reported | 56.9% |
| Terminal-Bench 4.0 (terminal coding) | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 Main | 46.4% | not reported | 42.4% | 52.1% |
| Chartography (reading charts) | 46.4% | 6.4% | 29.1% | 61.6% |
The jump from Haiku 4.5 is huge on every row. That is the headline, and it holds up.
Before you quote these numbers, know three things the table does not say:
- Every Haiku 5.5 score is at max effort. The system card confirms it. Max is the most expensive setting, and the default is medium. On OSWorld, medium scores 53.3%, not 72.4%.
- The OSWorld score is partial credit. The strict pass rate, where every step of the task has to be right, is 37.1% for Haiku 5.5, 48.8% for Sonnet 5.5 and 53.2% for Opus 5.5.
- Haiku 4.5's 0.0% on Terminal-Bench is real, but the setup differs. Haiku 4.5 has no effort setting, so Anthropic ran it with a fixed thinking budget. It passed none of its 660 trials.
Also, Anthropic ran GPT-6 Luna itself on OSWorld, through OpenAI's API. That is disclosed in a footnote. It is still one vendor running the other vendor's model.
Which effort level to use#
Haiku 5.5 is the first Haiku with an effort setting: low, medium, high, xhigh and max. This is where the money is, because the gap between settings is much bigger than the gap between models.
Anthropic charts score against cost per task at every effort level but does not print the numbers. They are in the page source, so I pulled them out. Pick a benchmark and an effort level:
Each dot is one effort setting, from low on the left to max on the right. Up and to the left is better: a higher score for less money per task.
Haiku 5.5 on medium: 53.3% for $0.13 a task.
- Sonnet 5.5 has no setting this cheap. Its cheapest is $0.68 for 57.9%.
- GPT-6 Luna for the same money or less: 37.5% on medium ($0.13).
- Haiku 4.5 has no setting this cheap. Its cheapest is $1.45 for 15.7%.
Extracted from the chart data embedded in Anthropic’s Claude Haiku 5.5 announcement. Costs are Anthropic’s own estimates at list prices, and Anthropic ran GPT-6 Luna itself on OSWorld, so treat this as a vendor chart. Haiku 4.5 has no effort setting and ran with a fixed thinking budget. “Same money” allows 5% over. Sonnet 5.5 uses its new $0.10 cache read price.
Here is what I measured from that data on October 7, 2026:
- Computer use is the clear win. Haiku 5.5 on medium scores 53.3% on OSWorld for $0.13 a task. GPT-6 Luna on max scores 48.9% for $0.21. That is a higher score for 39% less money.
- Against Haiku 4.5 it is not close. Haiku 5.5 on max scores 4.6 times higher on OSWorld (72.4% against 15.7%) and costs 58% less per task ($0.61 against $1.45).
- On office work, the same price per token does not mean the same price per task. Haiku 5.5 on high and GPT-6 Luna on max land in the same place on GDPval-AA: about 1420 to 1437 Elo for $0.09 a task. To reach its headline 1620, Haiku 5.5 on max spends $0.87 a task, about ten times Luna's most expensive setting.
- Max is where Sonnet 5.5 takes over for coding. On Terminal-Bench 4.0, Sonnet 5.5 on high scores 43% for $1.46. Haiku 5.5 on max scores 39.2% for $2.64.
- Max rarely earns its price. Going from xhigh to max on Haiku 5.5 adds 4.8 points on OSWorld for 2.2 times the cost, and 1.8 points on Humanity's Last Exam for 3.3 times the cost.
Anthropic's own chart shows the computer-use result:

And the coding result, where the Sonnet 5.5 line sits above the Haiku 5.5 line almost the whole way:

My rule, which matches Anthropic's effort guidance:
- Low for chat, classification, routing and short tool tasks.
- Medium as the default, including for subagents.
- High for knowledge work and longer agent tasks.
- Xhigh or max only after you have tested them against Sonnet 5.5 on your own work. Often Sonnet 5.5 on a lower setting is both better and cheaper.
I found the same pattern in Opus 5.5, where medium matched max at a fraction of the cost. If you want that data, it is in my Claude Opus 5.5 vs GPT-6 Sol and Luna benchmark breakdown.
When to use Haiku 5.5, and when not to#
Use it for:
- Summaries, compaction, classification, extraction and routing
- Live customer support and anything else where speed matters
- Browser and computer use tasks
- Subagents under a bigger model: Opus 5.5 or Sonnet 5.5 plans, Haiku 5.5 reads files and fetches facts
- High-volume document questions
Keep a bigger model for:
- The main coding agent on anything multi-step. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 against 39.2%, and 81.3% on SWE-Bench Pro against 64.8%.
- Work that keeps more than 100,000 tokens in context, where the price advantage shrinks.
- Anything where a wrong fact is expensive. More on that in the next section.
The customer quotes on the launch page point the same way. Asana measured up to 2.5 times faster inference per agent turn. AlphaSense, which makes about 8 million document-question calls a week, saw accuracy go from 0.76 to 0.84. Rogo uses it as the subagent that goes into a 10-K and pulls one number while a bigger model builds the deck. These are launch partners, so weigh them accordingly, but they all describe short, narrow, high-volume jobs.
What the system card says that the launch post does not#
Anthropic's system cards are more candid than its launch posts. This one runs 144 pages. Here is what stood out.
It refuses too much. In Anthropic's automated behavioral audit, Haiku 5.5 over-refused more than any other model tested, including Haiku 4.5. If you run it in a product, test your normal requests for false refusals.
It still makes things up at a Haiku rate. On a closed-book fact benchmark, Haiku 5.5 got more answers right than Haiku 4.5 but gave about as many wrong ones. It ranked behind every other Claude model in the comparison. Give it the documents. Do not ask it to remember.
It sometimes uses a leaked answer without saying so. In coding tasks where the solution was reachable inside the sandbox, Haiku 5.5 used it without telling the user 17% of the time. Haiku 4.5 did that 2% of the time. That is a step backward, and Anthropic says so.
It is bad at hiding things, which is good. When told to complete a harmful side task without a monitor noticing, Haiku 5.5 got away with it 6% of the time. Opus 5.5 managed 49%. Haiku 5.5 also owned up to concealed changes 96% of the time when asked for a summary, against 10% for Haiku 4.5.
It holds up against prompt injection. Anthropic calls it its most robust Haiku against injection attacks, close to its frontier models in coding and computer use tests. That matters for a model built to browse.
Safety filters can now stop a request. Haiku 5.5's cybersecurity safeguards are stricter than Haiku 4.5's. On Terminal-Bench, they flagged 12 of 660 trials, and every flagged trial failed because there is no fallback model. In your code, that shows up as stop_reason: "refusal".
If you are moving API code from Haiku 4.5, the migration guide lists the breaking changes. The ones most likely to bite: a budget_tokens thinking setting now returns an error, so do temperature, top_p and top_k, and so does prefilling the assistant's reply.
Two more things Anthropic shipped today#
Sonnet 5.5 got cheaper. Cache reads dropped from $0.20 to $0.10 per million tokens. Anthropic says that cuts the cost of most agent work on Sonnet 5.5 by about 20%, because agents reread the same context over and over.
Max and Team plans get monthly API credits. Rolling out this week: $100 a month on Max 5x, $200 on Max 20x, and up to $500 pooled on Team. The credits work on any model.
Do the math on that. $100 buys a billion input tokens of Haiku 5.5 at the launch price. If you have a Max plan and have been putting off building your own tool, the excuse is gone. I showed the whole path, from prompt to a live site on a real domain, in how to build and deploy a website with Claude Code. For hosting I use Hostinger, and the code MOE-LUEKER takes 10% off.
Quick answers#
Is Claude Haiku 5.5 free?#
Yes, in the Claude apps. Anthropic lists it for Free plan users on Claude.ai. The API is paid, from $0.10 per million input tokens.
What is the Claude Haiku 5.5 model ID?#
claude-haiku-5-5 on the Claude API, Google Cloud and Microsoft Foundry. On Amazon Bedrock it is anthropic.claude-haiku-5-5.
What is the context window?#
1 million tokens in, up to 128,000 tokens out. Prompts over 100,000 tokens are billed at the higher tier. The knowledge cutoff is June 2026.
Is Haiku 5.5 better than GPT-6 Luna?#
On Anthropic's numbers, yes on computer use, chart reading and terminal coding, and at the same price per token. On office work the two are close at the same cost per task. OpenAI has not published a head-to-head, so treat it as one side of the story.
Is Haiku 5.5 better than Sonnet 5.5?#
No. Sonnet 5.5 wins every row of the benchmark table. Haiku 5.5 is the cheaper and faster model for narrow jobs. On hard coding, Sonnet 5.5 is often cheaper per finished task too.
Does Claude Code have Haiku 5.5?#
Yes. Run /model and pick it. It runs at medium effort by default.
How does it compare to the bigger launches?#
It is a different kind of model. If you want the top end, read my Claude Opus 5.5 vs GPT-6 Astra comparison and the Opus 5.5 posts linked above.
The short version#
Claude Haiku 5.5 came out on October 7, 2026. It is available everywhere today, including the free Claude app. It costs $0.10 in and $0.50 out per million tokens until your prompt passes 100,000 tokens, and five times that after.
Run it on low or medium, keep its context small, and give it narrow jobs. Leave the hard coding to Sonnet 5.5 and Opus 5.5. And read the system card before you put it in front of customers, because it refuses more and still gets facts wrong more often than its bigger siblings.
I will test it in my own agent setup next and put the results on my YouTube channel.
One more thing#
Most of what a cheap model gets wrong, it gets wrong because the instructions were thin. The fix is to say more, and the fastest way to say more is out loud.
Rambleproof turns the way you actually talk, restarts and all, into clean text at your cursor. I use it for the long briefs I hand to coding agents. You can see what it does and what it costs first.

Published October 7, 2026, a few hours after the announcement. Sources: Anthropic's Claude Haiku 5.5 announcement, the Claude Haiku 5.5 system card, the model overview, what's new, pricing and effort docs. This post is not sponsored by Anthropic. I will update it as independent testing lands.
Moe shares tool walkthroughs and lessons from real projects. Mechanical engineer, then venture capital, now building AI tools for creators and small businesses. More about Moe