Claude Haiku 5.5 vs Haiku 4.5: What Ten Cents Now Buys

AIHubMix8 min read
Claude Haiku 5.5 vs Haiku 4.5: What Ten Cents Now Buys

Claude Haiku 4.5 scored 0.0% on Terminal-Bench 4.0. Claude Haiku 5.5 scores 39.2%, and it costs a tenth as much per token: $0.10 per million input tokens and $0.50 per million output tokens, against $1 and $5 for Haiku 4.5.

That is the upgrade in one line. The small Claude model went from "fine for classification, not for agents" to a usable subagent, and its price now matches OpenAI's GPT-6 Luna to the cent. Anthropic released it on October 7, 2026 as Claude Haiku 5.5, model ID claude-haiku-5-5, and it is already listed on AIHubMix.

The price comes with conditions, and so do the benchmark scores. This post covers what changed, how close Haiku 5.5 gets to Sonnet 5.5, where it is the wrong choice, and who should switch now.

What changed, and what didn't

Haiku 4.5 Haiku 5.5
Input / output price $1 / $5 $0.10 / $0.50
Context window 200K 1M
Max output 64K 128K
Thinking Manual budget, off by default Adaptive, on by default
Effort levels None Five, default medium
Tokens for the same text Baseline About 30% more
Browser use tool No Yes

Haiku 5.5 prices apply to prompts up to 100K tokens; longer prompts pay $0.50 / $2.50. Prices are per million tokens.

What stays the same: the job description. Anthropic still pitches Haiku at high-volume, cost-sensitive work (summaries, compaction, database queries, classification) and at subagent duty under a larger model. Anthropic says existing Haiku 4.5 prompts should work well without changes. Existing request code is a different matter: five request patterns that worked on Haiku 4.5 now return a 400 error, as Anthropic's migration guide lists and the migration post in this series walks through.

Two changes alter how the model behaves by default. Thinking is now on unless you switch it off, so a request that never mentioned thinking can come back with thinking blocks first and spend output tokens on them. And Haiku 5.5 is the first Haiku with an effort setting, which is now the main dial between cost and quality.

How it scores against Haiku 4.5, Luna, and Sonnet 5.5

Figures below come from Anthropic's launch materials and the Haiku 5.5 system card, as compiled by Kingy.ai. They are vendor-reported (Anthropic ran some of the GPT comparisons itself, such as OSWorld), and nobody has reproduced the full table independently yet.

Benchmark Haiku 5.5 Haiku 4.5 GPT-6 Luna Sonnet 5.5
GDPval-AA v2.1 (Elo) 1620 735 1437 1840
AA-Briefcase v1.1 1578 614 1336 1824
OSWorld 2.1 (offline) 72.4% 15.7% 48.9% 83.9%
Humanity's Last Exam 45.9% 10.2% n/a 56.9%
Terminal-Bench 4.0 39.2% 0.0% 16.4% 70.6%
FrontierCode 1.1 Main 46.4% n/a 42.4% 52.1%
Chartography 46.4% 6.4% 29.1% 61.6%

Humanity's Last Exam and Chartography are without tools. Sonnet 5.5's FrontierCode score is at xhigh effort.

Three readings matter more than any single row:

  • Against Haiku 4.5, it is a different class of model. Knowledge-work ratings roughly double, and the agentic rows go from near zero to usable. Haiku 4.5 could not finish a Terminal-Bench task; Haiku 5.5 finishes about two in five.
  • Against GPT-6 Luna, it leads on every shared row at the same list price. The widest gaps are on computer use (72.4% vs 48.9%) and Terminal-Bench (39.2% vs 16.4%).
  • Against Sonnet 5.5, the gap depends on the task. On FrontierCode, Haiku is within six points. On Terminal-Bench, Sonnet scores 70.6% to Haiku's 39.2%. Anthropic itself says Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding.

Read the fine print before quoting these numbers

The headline scores were run at max effort, the most expensive setting. The default is medium, and the system card's own medium-effort runs come in lower: 1277 on GDPval-AA instead of 1620, and 1372 on AA-Briefcase instead of 1578. If you deploy at the default, compare against those numbers, not the launch chart.

The OSWorld figure needs a second look too. The 72.4% is partial credit, averaged over checkpoints. The share of tasks where Haiku 5.5 completed every checkpoint is 37.1% (Luna: 17.1%). For a browser agent that must finish the whole job, 37.1% is the number to plan around.

What early customers reported

Anthropic's launch post quotes several customers who tested the model before release. These are vendor-selected, but they point at the same workloads:

  • AlphaSense runs about 8 million "Ask in Document" calls a week. On 400 test queries, Haiku 5.5 scored 0.84 against Haiku 4.5's 0.76.
  • Box saw an 11-point gain over Haiku 4.5 at about half the latency.
  • Asana measured over 30% lower latency per task and up to 2.5x faster inference per agent turn, compared with its current model.
  • HubSpot got 92.8% on its CRM suite, the best score it had seen on that suite.
  • Cognition uses Haiku 5.5 as a "sidekick" under Opus 5.5 in Devin Fusion and reports that the pair holds a FrontierCode score of 66.2 at lower cost and latency.

The pattern: short, repeated work over documents, and subagent duty under a bigger model.

Where Haiku 5.5 is the wrong choice

  • Complex agentic coding. Terminal-Bench is where Sonnet 5.5 pulls furthest ahead. If a task needs many steps, ambiguous requirements, or a long chain of edits, start with Sonnet 5.5 or Opus 5.5.
  • Prompts over 100K tokens. The 1M window is real, but the price is not flat across it. Past 100K tokens the whole request moves to $0.50 / $2.50, and because the new tokenizer counts about 30% more tokens, that line sits near 77K tokens as Haiku 4.5 counted them. The pricing post in this series works through the numbers.
  • Work that needs xhigh or max to pass. Output gets long and expensive at the top settings. Anthropic's own guidance is to compare against Sonnet 5.5 before settling there.
  • Teams on Priority Tier. Haiku 5.5 does not support it, so a Haiku 4.5 Priority Tier commitment does not carry over.
  • Security testing. Haiku 5.5's cyber safeguards are stricter than Haiku 4.5's and block penetration testing. Expect refusal stop reasons on that kind of work, with no server-side fallback to another model.

Try it through AIHubMix

Haiku 5.5's effort setting lives in the Messages API's output_config, so the Claude native endpoint on AIHubMix is the most direct route. Install a current Anthropic SDK (pip install -U anthropic) and point it at AIHubMix:

import os
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com",
)

response = client.messages.create(
    model="claude-haiku-5-5",
    max_tokens=2000,  # leave room for thinking, not just the answer
    output_config={"effort": "low"},
    messages=[{
        "role": "user",
        "content": "Classify this ticket as billing, bug, or other: "
                   "'I was charged twice for October.'",
    }],
)

# Thinking blocks can come first, so pick the text block by type.
answer = next(block.text for block in response.content if block.type == "text")
print(answer)
print(response.usage)

One thing to check on your first run: read response.usage.output_tokens at low and again at high on the same prompt: if the numbers don't move, the effort setting is not reaching the model. Haiku 5.5 runs with the full 1M context window on AIHubMix, so long prompts work as-is; just keep the 100K price line in mind.

Who should switch, who should wait

Switch now if you run classification, extraction, routing, summaries, or document Q&A with prompts well under 100K tokens. You get a much stronger model at a fraction of the token price. Plan an hour for the request-code changes, then test effort low against medium on your own evals.

Switch now if you use a large model for subagent work it is overqualified for: pulling a figure from a filing, checking a file, summarizing a tool result. This is the role Anthropic and its launch customers emphasize.

Wait a sprint if your integration uses computer use with the computer_20250124 tool, prefill, or temperature=0. Each of those returns a 400 on Haiku 5.5, and fixing them is code work, not a model-string swap.

Stay where you are if your workload is long-prompt (regularly over about 77K Haiku 4.5 tokens), depends on Priority Tier, or is complex agentic coding. For the last group, the useful comparison is Haiku 5.5 against Sonnet 5.5 on cost per finished task, not Haiku 5.5 against Haiku 4.5.

When you're ready to compare, the AIHubMix model list puts Haiku 5.5, Sonnet 5.5, and Opus 5.5 behind one API key, so the same eval can run against all three.

FAQ

Is Claude Haiku 5.5 better than Haiku 4.5 at everything?
On every benchmark in Anthropic's launch table, yes, often by a wide margin. The exceptions are practical rather than quality: Priority Tier is not supported, prompts over 100K tokens cost more per token than the short-prompt rate, and several old request patterns now return errors.

Is Haiku 5.5 really 90% cheaper?
The per-token price is 90% lower for prompts up to 100K tokens and 50% lower above that. Weighting by the request mix (about 90% of Haiku 4.5 requests were under 100K tokens) and the new tokenizer, which counts about 30% more tokens for the same text, Anthropic estimates an average saving of around 75%. Default-on thinking adds output tokens on top of that, so your actual bill also depends on the effort setting.

How does Haiku 5.5 compare with GPT-6 Luna?
Both list at $0.10 per million input tokens and $0.50 per million output tokens. In Anthropic's reported results, Haiku 5.5 scores higher on every benchmark where both appear, most clearly on computer use and Terminal-Bench. Luna keeps its base price up to a longer prompt length, so long-prompt workloads should be priced separately.

Can Haiku 5.5 replace Sonnet 5.5 for coding?
For narrow, well-scoped changes and subagent tasks, it can be worth testing. For complex multi-step agentic coding, Sonnet 5.5 scores far higher on Terminal-Bench 4.0 (70.6% vs 39.2%), and Anthropic recommends Sonnet or Opus for that work.

Are the benchmark scores at the default setting?
No. The headline scores were run at max effort. The default is medium, where the system card reports lower results, for example 1277 instead of 1620 on GDPval-AA.

Does the 1M context window mean I can send 1M-token prompts cheaply?
You can send them, but anything over 100K tokens is billed at $0.50 input and $2.50 output per million tokens for the whole request. AIHubMix supports the full 1M window as well.

What do I have to change in my code to switch?
At minimum: the model ID, any manual thinking budget, any temperature or top_p setting, any assistant prefill, and code that reads the first content block as the answer. The migration post in this series has the full list.

Keep reading: the Claude Haiku 5.5 series

Sources