The solution
Launching Claude Haiku 5.5 on October 7, 2026, Anthropic cut token prices by 90% for requests below 100,000 tokens — $0.10 per million input and $0.50 per million output, exactly matching OpenAI's GPT-6 Luna — and positioned the model as the supporting worker for its pricier Opus and Sonnet: a larger model might assemble a financial presentation while Haiku retrieves one revenue figure, as customer Rogo's example put it. About 90% of Haiku 4.5 requests fall under the threshold, and workloads were estimated to cost about 75% less once request sizes and a less generous tokenizer were counted.
The structure is the strategy: above 100,000 tokens, prices step up five-fold to $0.50/$2.50 — still a 50% cut, but a real premium for long context, where Luna's own surcharge only begins at 272,000 input tokens. Simon Willison, testing the model on launch day, showed his long-prompt token counter reading about 1.25x Haiku 4.5's tokens — 'a hidden price increase' — and noted that beyond 100K tokens 'Luna looks like a much better deal.' The tiering prices the agent economy's repetitive calls to win it, while long-context work pays.
The quieter move was the subscription bridge: monthly API credits for Max and Team subscribers — $100 for Max 5x, $200 for Max 20x, up to $500 pooled for Team. Willison observed the credits exactly match the subscription's own cost, turning the plan into prepaid API budget — use them or lose them monthly, with auto-reload optional so balance exhaustion stops requests instead of surprising the card.
Why it worked
Agent architectures route most calls to the small model, so the under-100K tier is where platform loyalty is won; pricing it at parity with Luna removes the last reason to switch.
The 5x long-context step keeps margin on the minority of requests where bigger models are already competitive.
Subscription-redeemable API credits pull consumer subscribers into the developer platform without discounting list prices — and their monthly expiry manufactures renewal.
What can be applied
In multi-model systems the cheap model does most of the calls — price that tier to win the whole agent stack, and let the long tail pay the margin.
Aftermath
At launch Anthropic also halved Sonnet 5.5 cache-read prices ($0.20 to $0.10 per million) and reported vendor-run benchmark leads over Luna on GDPval-AA, OSWorld and Terminal-Bench, with adjustable effort levels defaulting to medium; independent verification remained pending.
FOLLOW THE EVIDENCE
The sources
- Anthropic launches Claude Haiku 5.5 with 90% API price reduction, matching GPT-6 Luna venturebeat.com
- Introducing Claude Haiku 5.5 simonwillison.net