Claude Haiku 5.5, launched on 7 October 2026, is priced in two tiers. Per Let’s Data Science, VKTR and Computing for Geeks, prompts up to 100,000 tokens cost $0.10 per million input tokens and $0.50 per million output tokens; prompts above 100,000 cost $0.50 and $2.50. Haiku 4.5 cost $1 and $5. Against that baseline the short-prompt tier is 90% cheaper and the long-prompt tier is 50% cheaper. VKTR quotes Anthropic as saying the model costs around 75% less to run on average. Three numbers, one price list. This piece works out which applies to which workload, and adds two things the launch coverage mostly leaves out: a new tokenizer, and the way the 100,000-token threshold behaves.

A note on sourcing. We could not read Anthropic’s own announcement or pricing page, so every price, benchmark and date below is as relayed by the three outlets named above. They agree with each other on the prices. Where a figure comes from Anthropic’s own benchmarking, we say so, because none of it has been independently reproduced.

The price list, tier by tier

  • Haiku 4.5: $1.00 input and $5.00 output per million tokens.
  • Haiku 5.5, prompts up to 100,000 tokens: $0.10 input, $0.50 output. Discount against Haiku 4.5: 90% on both.
  • Haiku 5.5, prompts over 100,000 tokens: $0.50 input, $2.50 output. Discount: 50% on both.
  • Batch processing keeps its 50% discount on top (VKTR); Computing for Geeks lists batch rates of $0.05/$0.25 and $0.25/$1.25 for the two tiers.

Computing for Geeks adds a detail that matters for budgeting: the long-prompt rate applies to every token in the request, including output and cached tokens, not only the tokens above 100,000. That makes the threshold a cliff rather than a slope. Take a prompt of exactly 100,000 tokens with a 1,000-token reply. At the short tier it costs 0.1 million × $0.10 = $0.01 for input plus 0.001 million × $0.50 = $0.0005 for output, so $0.0105. Add one token to the prompt and the same request is billed at the long tier: $0.05 plus $0.0025, so $0.0525, five times as much. On Haiku 4.5 the same request cost $0.10 plus $0.005, or $0.105. The saving drops from 90% to 50% across a single token. That arithmetic is ours.

Why the tokenizer matters more than the headline

Both VKTR and Computing for Geeks report that Haiku 5.5 uses a new tokenizer that produces about 30% more tokens for the same text. Neither gives a measured range, so treat 30% as an average, not a guarantee. If it holds, a customer’s bill is the per-token price multiplied by roughly 1.3, and the effective discounts shrink:

  • Short tier: $0.10 × 1.3 = $0.13 per million old-tokenizer-equivalent tokens, so about 87% cheaper than Haiku 4.5, not 90%.
  • Long tier: $0.50 × 1.3 = $0.65, so about 35% cheaper, not 50%.

The same factor moves the cliff. A prompt that Haiku 4.5’s tokenizer counted as 100,000 tokens would count as about 130,000 under the new one. Put the other way, the 100,000-token threshold sits at roughly 77,000 tokens in the old counting (100,000 ÷ 1.3 = 76,923). Teams migrating long-context pipelines whose prompts used to land between about 77,000 and 100,000 tokens would find themselves in the expensive tier without having changed anything. Computing for Geeks makes the same qualitative point, that prompts near 100,000 tokens may cross into the higher tier after recounting; the 77,000 figure is our arithmetic on its 30% number.

One price list, three headline numbers

The 90%, 75% and 50% figures in circulation are all describing the same price list from different positions. A blended saving is the sum, across tiers, of each tier’s share of your current Haiku 4.5 spend multiplied by that tier’s discount. The share has to be measured in spend, not request count, because long requests cost far more than short ones. With that definition, our calculations are:

  • A 90/10 split of Haiku 4.5 spend between short and long prompts, list prices only: 0.9 × 0.10 + 0.1 × 0.50 = 0.14 of the old bill, a saving of 86%.
  • The same split with the 30% tokenizer effect: 0.9 × 0.13 + 0.1 × 0.65 = 0.182, a saving of about 82%.
  • What mix produces Anthropic’s 75% average? With the tokenizer effect, 0.13s + 0.65(1 − s) = 0.25 gives s ≈ 0.77: roughly 77% of old spend on short prompts and 23% on long ones. Without the tokenizer effect, s = 0.625.

So the 75% claim is consistent with the published rates. It follows from a traffic mix in which about a quarter of spend goes to prompts over 100,000 tokens, or from a heavier long-context share if the tokenizer is ignored. What it is not is a number any individual team should assume. A support-bot team with short prompts could land near 87%; a team that stuffs whole codebases into context could land near 35%. The only way to know is to bucket last month’s Haiku 4.5 spend by prompt size and apply the two tiers. Other roundups quote 90% under 100,000 tokens and 50% above, and VKTR estimates 85% to 90% for short-prompt workloads after allowing for the tokenizer. These are the same table read at different points.

Price per token is not cost per task

Haiku 5.5 adds an effort parameter with five levels (low, medium, high, xhigh, max), and adaptive thinking is on by default at medium effort, per VKTR and Computing for Geeks. Haiku 4.5 had no effort setting. Thinking tokens count toward the output limit and are billed as output, so a model that reasons longer can cost more per finished task even at a lower per-token price.

We can bound how much longer. At the short tier, output is 10 times cheaper per token than Haiku 4.5 ($0.50 versus $5), but each unit of text costs 1.3 times as many tokens, so Haiku 5.5 would have to produce more than about 7.7 times the output tokens its tokenizer already implies (10 ÷ 1.3) before a task cost more than it did on Haiku 4.5. At the long tier the output price ratio is only 2 ($2.50 versus $5), and the break-even falls to about 1.5 times the tokens (2 ÷ 1.3). That is a much smaller margin for reasoning to use up. Computing for Geeks reports one test in which max effort used all 32,000 allowed output tokens on a single task. That is one reviewer’s run and not a typical figure, but at $0.50 per million, 32,000 tokens cost $0.016, so the short tier absorbs even a heavy reasoning budget. The risk is concentrated in long-context, high-effort jobs.

Computing for Geeks also reports Anthropic’s GDPval-AA scores falling from an Elo of 1620 at max effort to 1277 at medium, against 735 for Haiku 4.5. If that is right, the effort level a team chooses changes quality by a margin comparable to the gap between model generations, so effort belongs in any cost comparison. Both numbers are vendor-reported.

What the benchmarks do and do not show

The headline capability number is OSWorld 2.1, an agentic computer-use benchmark, scored on an offline subset: Haiku 5.5 at 72.4% against 15.7% for Haiku 4.5, as reported by Let’s Data Science from Anthropic’s materials. Computing for Geeks relays Anthropic’s comparison figures of 48.9% for OpenAI’s GPT-6 Luna and 83.9% for Anthropic’s own Sonnet 5.5, run at max effort. All of these are Anthropic’s figures, on a subset, with computer-use support in the SDKs still in beta. For Terminal-Bench 4.0, Computing for Geeks lists 39.2% from Anthropic and 33% measured by Artificial Analysis in its own harness. That is the only independent number in the coverage we reviewed; it is lower, though harness differences could explain some of the gap. Anthropic did not publish SWE-bench Verified, GPQA Diamond, AIME or tau-bench results for this model, according to the same source.

The context window is reported at 1 million tokens, up from 200,000, with up to 128,000 output tokens, by both VKTR and Computing for Geeks; Let’s Data Science does not state it. Per VKTR, Anthropic has committed to keeping the model available until at least 7 October 2027, and it launched on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. A 1 million-token window sits awkwardly beside a 100,000-token pricing cliff: the model can take far larger inputs than the cheap tier covers, so most of the long-context capability is priced at the 50% discount, or about 35% after the tokenizer.

What the evidence does not establish

None of the coverage includes an independent cost-per-task measurement for Haiku 5.5, a measured tokenizer ratio, or a reproduction of the OSWorld 2.1 subset result. The 30% tokenizer figure is the weakest link in our arithmetic, since both outlets state it without a source or a range, and every adjusted number above scales with it. If the true factor were 1.1 or 1.5, the short-tier discount would be 89% or 85%, and the long-tier discount 45% or 25%. The practical test is straightforward: replay a day of real prompts through both models, count billed tokens, and compare the invoices at the effort level you intend to run.