Migrating from Claude Haiku 4.5 to 5.5: Five 400 Errors and the Quiet Changes

AIHubMix9 min read
Migrating from Claude Haiku 4.5 to 5.5: Five 400 Errors and the Quiet Changes

Changing claude-haiku-4-5 to claude-haiku-5-5 is the smallest part of this migration. Five request patterns that worked on Haiku 4.5 now return a 400 error, and several more changes fail no request but alter what you get back, what it costs, or how the model behaves inside an agent.

Anthropic says existing Haiku 4.5 prompts should work well on Haiku 5.5 without changes. The request code around those prompts is a different story. This post lists each problem as you'll meet it: what you'll see, why it happens, and how to fix it, followed by a checklist. The authoritative reference is Anthropic's Haiku 5.5 migration guide.

Triage: match the symptom

What you see Cause Fix
400 on a request with a thinking budget Manual thinking removed Adaptive thinking plus effort
400 with temperature, top_p, or top_k Sampling parameters locked Remove them
400 when messages end on an assistant turn Prefill removed End on a user turn
400 on computer use Old computer tool rejected Move to the computer toolset
400 after editing earlier turns Thinking bound to history Keep history append-only
Parser returns empty or wrong text Thinking block comes first Select blocks by type
Reply cut off or missing Thinking counts toward the cap Raise max_tokens or lower effort
Token counts and bills up about 30% New tokenizer Recount on the new model
Response with stop reason refusal New safety classifiers Handle it in your client

The first five fail loudly. The rest fail quietly, which makes them more expensive to find.

The five loud failures

1. Manual thinking budgets

What you'll see: a 400 on any request that sends thinking: {"type": "enabled", "budget_tokens": N}.

Why: Haiku 4.5 only supported manual extended thinking with a token budget. Haiku 5.5 supports adaptive thinking only, and controls depth with effort.

Fix: send {"type": "adaptive"} or leave thinking out, and pick an effort level. Where the old budget was small to save tokens, choose a low level.

# Before: Haiku 4.5
thinking={"type": "enabled", "budget_tokens": 8000}

# After: Haiku 5.5
thinking={"type": "adaptive"},
output_config={"effort": "medium"},

2. Sampling parameters

What you'll see: a 400 when a request sets temperature, top_p, or top_k.

Why: Haiku 5.5 accepts only the defaults: temperature of 1 and top_p of 0.99. Any other value of either, any top_k, or sending both temperature and top_p returns a 400, whether or not thinking is used. A top_p of 1 is rejected too.

Fix: remove all three. The common case is temperature=0 on a classifier, used to get stable labels. Replace it with structured output or a tool whose input is an enum, so the label set is enforced by schema rather than by sampling. Also check SDK wrappers and gateways that add default sampling values on your behalf.

3. Assistant prefill

What you'll see: a 400 when the last entry in messages is an assistant turn, even with thinking off.

Why: prefill is not supported on Haiku 5.5, matching the rest of the current Claude lineup.

Fix: end messages with a user turn, and replace the prefill by what it was for. Format control becomes structured output (output_config.format). A prefilled preamble becomes a system-prompt instruction to answer directly. A continuation of an interrupted reply moves into the user message: "Your previous response ended with [text]. Continue from there."

4. Computer use

What you'll see: a 400 on the Claude API or Google Cloud when the request declares the computer_20250124 tool.

Why: on those platforms Haiku 5.5 supports computer use only through the newer toolset, computer_toolset_20260801.

Fix: drop the computer-use-2025-01-24 beta header, replace the tool entry with {"type": "computer_toolset_20260801"}, and update the agent loop: dispatch on each tool_use block's name and toolset_name rather than on input.action, handle every such block in a turn, and echo toolset_name on results. Zoom is on by default; if your environment doesn't implement it, disable it in the toolset config. On Amazon Bedrock, check the computer use tool's compatibility notes before choosing a version. The same toolset family also brings browser use, which Haiku 4.5 never had.

5. Editing earlier turns

What you'll see: a 400 when a request sends back a thinking block after something before it changed: the system prompt, the tool list, or an earlier message.

Why: a Haiku 5.5 thinking block stays valid only while everything sent before it is unchanged. The check is enforced by default for accounts created on or after August 31, 2026, and on older accounts only when a request opts in.

Fix: keep conversations append-only. Common culprits are a system prompt with a timestamp, a tool list that grows when a plugin connects, client-side truncation, and reminders injected into history and stripped on the next turn. For per-turn instructions, Haiku 5.5 supports system messages inside messages, with no beta header, which add context without editing what came before.

The quiet failures

Thinking blocks come first. Thinking is on by default, so a response can start with one or more thinking blocks. Code that reads response.content[0].text as the answer breaks or returns empty text. Select blocks by type.

Thinking text is empty by default. Haiku 4.5 returned summarized thinking. Haiku 5.5 returns thinking blocks with an empty text field and only a signature. If your UI showed reasoning summaries, set thinking: {"type": "adaptive", "display": "summarized"}. Either way, pass thinking blocks back unchanged with tool results; a serializer that drops empty blocks removes them.

max_tokens now has to cover thinking. A cap sized for a short answer can be used up by thinking, ending the response with stop_reason: "max_tokens" before any text. Raise the cap or lower effort.

The same text is about 30% more tokens. The new tokenizer changes usage fields, count_tokens results, context budgets, and any max_tokens tuned for Haiku 4.5. It also moves the 100K-token price line to about 77K tokens as Haiku 4.5 counted them. Recount real prompts with the model set to claude-haiku-5-5 before trusting a cost dashboard.

The default effort is medium. Haiku 4.5 had no effort setting. Haiku 5.5 defaults to medium, which may be more thinking than a simple route needs. Set it explicitly.

Thinking blocks stay with the account that made them. If your service replays stored conversations through a different API account, Haiku 5.5's thinking blocks are silently dropped and the request runs without that reasoning. Replay each conversation through the account that produced it.

Priority Tier doesn't carry over. Haiku 5.5 doesn't support Priority Tier, so plan capacity separately if you rely on it for Haiku 4.5.

Gateway listings can differ. The Haiku 5.5 page on AIHubMix currently lists a 200K context length, while Anthropic specifies 1M. Confirm the limit on the route you use before migrating long-prompt workloads.

Behaviour changes that matter for agents with real permissions

Refusals are new, and nothing catches them for you. Haiku 5.5 runs safety classifiers in four categories: cyber, bio, frontier LLM development, and general harms. A decline comes back as a normal HTTP 200 with stop_reason: "refusal" and a category in stop_details. Unlike Sonnet 5.5 and Opus 5.5, Haiku 5.5 has no server-side fallback: a list of fallback models returns a 400, and the default fallback mode leaves the request declined. Check stop_reason before reading content, and decide in your own code whether to rephrase, escalate to a larger model, or stop. Per the launch post, the cyber safeguards allow a wider range of defensive work than Sonnet 5.5's but block penetration testing.

User text inside tool results may be ignored. Haiku 5.5 is trained to resist prompt injection through tool results. If your harness delivers a message the user typed mid-task inside a tool_result block, the model can treat it as untrusted and ignore it. Put mid-turn user input in a text block after the last tool result, and keep harness notices in a separate system message.

At low effort, agents can stop early or skip checks. With a long coding-agent system prompt at low, Haiku 5.5 sometimes hands the task back before it is done, and at low and medium it sometimes reports a code change as done without running a test. Anthropic's Haiku 5.5 prompting guide has short instructions for both. For an agent that can write files or run commands, an unverified "done" is the more dangerous of the two.

Forcing a tool skips thinking. Forced tool_choice is still accepted, but the model then calls the tool without thinking first. For tools with side effects, auto plus a clear instruction lets the model reason before acting.

Search tools need today's date. When Haiku 5.5 has a search tool, give it the current date in the system prompt or the tool description. In Anthropic's testing this grounded answers in recent results.

A migrated request through AIHubMix

A Haiku 4.5 classifier that used temperature=0, a thinking budget, and a prefilled { for JSON, rewritten for Haiku 5.5 on the AIHubMix Claude native endpoint:

import os
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com",
)

r = client.messages.create(
    model="claude-haiku-5-5",
    max_tokens=2000,                     # room for thinking plus the JSON
    output_config={
        "effort": "low",                 # replaces the old thinking budget
        "format": {                      # replaces the prefill and temperature=0
            "type": "json_schema",
            "schema": {
                "type": "object",
                "properties": {
                    "label": {"type": "string", "enum": ["billing", "bug", "other"]}
                },
                "required": ["label"],
                "additionalProperties": False,
            },
        },
    },
    messages=[{"role": "user", "content": "Ticket: 'I was charged twice for October.'"}],
)

if r.stop_reason == "refusal":
    raise RuntimeError(f"declined: {r.stop_details}")
text = next(b.text for b in r.content if b.type == "text")
print(text)

Whether a gateway forwards output_config fields and newer beta headers unchanged is worth confirming on your first test run. When the migrated route passes your evals, the AIHubMix model list makes it easy to point the same code at Sonnet 5.5 for any task type that keeps failing on Haiku.

Migration checklist

  1. Change the model ID to claude-haiku-5-5, with no date suffix.
  2. Replace every thinking budget with adaptive thinking and an explicit effort level.
  3. Remove temperature, top_p, and top_k, including defaults added by wrappers.
  4. Replace assistant prefills with structured output, system instructions, or user-turn continuations.
  5. Move computer use to the computer toolset and update the agent loop.
  6. Make conversation history append-only if thinking blocks are replayed.
  7. Read response content by block type, and keep empty thinking blocks when replaying.
  8. Raise max_tokens on short-answer routes, or lower effort.
  9. Handle the refusal stop reason before reading content; don't configure server-side fallbacks.
  10. Recount prompt tokens on the new model and re-baseline cost dashboards.
  11. Check which prompts now cross 100K tokens and trim or split them.
  12. Set display to summarized if users saw reasoning summaries.
  13. Deliver mid-turn user input outside tool results.
  14. Give search-enabled agents today's date.
  15. Re-check rate limits, Priority Tier needs, and your gateway's context limit before moving volume.

FAQ

Will my Haiku 4.5 prompts work on Haiku 5.5?
Anthropic says existing prompts should work well without changes. The request parameters around them are what break: thinking budgets, sampling settings, prefills, and the old computer use tool all return errors.

Why does my classifier fail now that I removed temperature 0?
It shouldn't fail, but labels may vary more. Use structured output or a tool with an enum field so the allowed labels are enforced by schema. That is more reliable than temperature 0 ever was.

Can I still turn thinking off?
Yes, at low, medium, and high effort. At xhigh and max, disabling thinking returns an error. Anthropic recommends a lower effort level instead, because the model can skip thinking on simple requests by itself.

What should my code do when Haiku 5.5 refuses?
Check the stop reason before reading the content. Haiku 5.5 has no server-side fallback, so your code decides whether to rephrase, send the request to a larger model, or return an error to the user.

Why did my token usage go up after migrating?
Two reasons. The new tokenizer counts about 30% more tokens for the same text, and thinking is on by default, adding output tokens. Lower effort and recount your prompts on the new model.

Do I need to change anything for prompt caching?
Usually not, and it gets easier: the minimum cacheable prompt drops from 4,096 to 512 tokens, and thinking blocks from earlier turns stay in the cached prefix by default. Avoid editing earlier turns, which now invalidates thinking blocks as well as the cache.

Can a conversation move from Haiku 5.5 up to a larger model?
Yes. Sonnet 5.5 and Opus 5.5 read Haiku 5.5's thinking blocks, so a conversation escalated to either keeps its earlier reasoning. For other target models, check the preserved thinking documentation first.

Keep reading: the Claude Haiku 5.5 series

Sources