Claude Haiku 5.5: migration notes and pricing analysis
Anthropic has released Claude Haiku 5.5. The model is built for high-volume, cost-sensitive work such as summaries, classification and database queries. It costs around 75% less than Haiku 4.5 on average, but a new tokenizer and a 100,000-token price threshold mean the real cost depends on your workload. This post collects the prices, the benchmark results and what to check when you move from Haiku 4.5. Most of the figures come from Anthropic's announcement; the tokenizer and timing measurements come from Simon Willison's review.
Pricing
Haiku 5.5 is priced in two tiers by prompt length. Up to 100,000 tokens, input costs $0.10 and output $0.50 per million tokens. Above that, the price goes up five times. According to Anthropic, around 90% of requests to the previous Haiku already stayed under 100,000 tokens.
| Model | Input | Output |
|---|---|---|
| Haiku 5.5 (≤100k tokens) | $0.10 | $0.50 |
| Haiku 5.5 (>100k tokens) | $0.50 | $2.50 |
| Haiku 4.5 | $1.00 | $5.00 |
| GPT-6 Luna (≤272k tokens) | $0.10 | $0.50 |
| GPT-6 Luna (>272k tokens) | $0.20 | $0.75 |
Prices are per million tokens.
The 100,000-token threshold
Below 100,000 tokens, Haiku 5.5 costs exactly the same as OpenAI's GPT-6 Luna. The difference starts above the threshold. Haiku 5.5's price jumps fivefold at 100,000 tokens, while Luna keeps the same price up to 272,000 tokens. After that it only rises to $0.20 / $0.75. That makes Luna clearly cheaper for workloads with long context.
The tokenizer effect
Haiku 5.5 uses a new, less efficient tokenizer. The tokenizer is the layer that splits text into the pieces the model counts, the tokens. The more tokens the same text is split into, the more you pay. Simon Willison measured this with his Claude Token Counter tool: the same long prompt takes around 1.25 times as many tokens on Haiku 5.5 as on Haiku 4.5.
So for the same text you pay around 25% more than the token price suggests. Next to the 75% price cut, that difference is small. It does matter close to the threshold: a prompt that takes 85,000 tokens on Haiku 4.5 grows to around 106,000 on Haiku 5.5 and lands in the expensive tier.
Performance
In the benchmarks Anthropic published, Haiku 5.5 beats Haiku 4.5 by a wide margin. It is ahead of GPT-6 Luna on all three tests.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna |
|---|---|---|---|
| GDPval-AA v2.1 | 1620 | 735 | 1437 |
| OSWorld 2.1 | 72.4% | 15.7% | 48.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% |
GDPval-AA scores performance on knowledge work. OSWorld measures how far an agent gets with long, multi-step tasks on a real computer; the figures are for the test's offline subset. Terminal-Bench measures agentic coding tasks.
The effort setting
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting. The effort setting decides how much the model reasons before it answers. A low setting is faster and cheaper, a high one slower but more accurate. Reasoning cannot be switched off entirely, and the default is medium.
The setting has a large effect on time. In Simon Willison's test, an SVG drawing at low took 7 seconds and cost 0.0936 cents. The same request at max took 5 minutes 9 seconds and cost 3.38 cents. On speed, Anthropic's announcement quotes an engineer at Asana: in their agent product, task completion latency dropped by more than 30%, and inference per agent turn got up to 2.5 times faster.
Migration notes
The first thing to change when moving from Haiku 4.5 is the model name: claude-haiku-5-5. Besides the Claude Platform, the model is available on AWS, Google Cloud and Microsoft Azure. Anthropic's migration guide lists the details.
It is worth redoing the cost calculation before you switch, because the tokenizer changed. Measure the token counts of your current prompts again with Haiku 5.5 and check the requests that come close to the 100,000 threshold separately. Since reasoning cannot be switched off, try low for latency-sensitive work. The default medium can be slower than such work needs.
Other updates
On the same day, Anthropic halved the price of cache reads for Claude Sonnet 5.5: $0.10 per million tokens instead of $0.20. According to Anthropic, this makes Sonnet 5.5 around 20% cheaper on most agentic work.
Max and Team subscribers now get monthly API credits to use on the Claude Platform. That is $100 for Max 5x, $200 for Max 20x and up to $500 for Team, pooled across users. The credits do not roll over. To claim them, pick the API organization that should receive them under Settings → Billing. If you turn off auto-reload for the API, requests stop when the balance runs out. That is the easiest way to avoid a surprise bill while experimenting with the credits.
The Python and TypeScript SDKs also gained beta support for computer use and browser use.
Which one to pick
| Scenario | First choice | Alternative |
|---|---|---|
| High volume, under 100k tokens | Haiku 5.5 | GPT-6 Luna (same price) |
| Prompts regularly over 100k tokens | GPT-6 Luna | Haiku 5.5, in its expensive tier |
| Currently on Haiku 4.5 | Haiku 5.5, after re-measuring tokens | Stay on Haiku 4.5 |
| Latency-sensitive agent steps | Haiku 5.5, low |
Haiku 5.5, medium |