Lower prices for a smaller-task model
Anthropic released Claude Haiku 5.5 on October 7, targeting work such as classification, summaries and support tasks. The company's listed first-party API rates start at $0.10 per million input tokens and $0.50 per million output tokens. Haiku 4.5's corresponding base rates are $1 and $5.
That is a 90% reduction in those unit prices for prompts up to 100,000 input tokens. It is not a promise that every task will cost 90% less. A token is a chunk of text counted by the service, and changing the model can change the number of chunks used.
The threshold and the tokenizer matter
For prompts above 100,000 input tokens, Haiku 5.5's listed rates rise to $0.50 for input and $2.50 for output, per million tokens. Both sides of the bill therefore need attention when working with long documents.
The migration guide says the new tokenizer produces approximately 30% more tokens for the same text than Haiku 4.5, with the difference depending on the content. Existing token estimates and output limits need checking.
Anthropic estimates around 75% lower total cost on average in its announcement. That is the company's calculation, not an independent result for a reader's workload. Listed prices also do not remove the need to check a cloud provider's own charges.
Some old settings do not carry across
Haiku 5.5 uses adaptive thinking by default: the model can decide how much reasoning to apply. The documentation says thinking tokens count toward the output limit. Lower effort can suit simpler work, but the appropriate setting depends on the task.
The migration guide warns that legacy reasoning-budget settings and certain sampling settings can produce errors. Structured outputs are supported on the Claude API and Google's Vertex AI, but are not available for Haiku 5.5 on Amazon Bedrock. Moving a model name is not necessarily a complete migration.
Compare the cost of a successful task
The practical comparison is a set of real tasks with the same acceptance standard: did the answer work, how long did it take and what was charged? Count failed responses and any extra calls needed to correct them, not only the price of one response.
A small model can be economical when a task is narrow and easy to verify. Lower unit prices do not establish that it is the right choice for every demanding workflow. The release makes a new option available; a workload test determines whether that option is useful.
Sources
- Anthropic: Claude Haiku 5.5 announcement, October 7Primary company announcement supports release date, intended narrow tasks, token rates and attributed around-75% average total-cost estimate. No independent benchmark superiority, guaranteed savings or safety validation claimed.
- Claude Platform: pricing documentationPrimary table verifies USD per-million-token input/output rates, Haiku 4.5 comparator, up-to-100,000 versus over-100,000 input-token bands and separate cloud-provider pricing. The 90% unit-rate reduction is arithmetic, not a workload result.
- Claude Platform: migrating to Claude Haiku 5.5Primary guide supports approximate content-dependent 30% tokenizer increase, default adaptive thinking, output-token accounting, legacy-setting errors and Bedrock structured-output limitation. No SDK or customer system was changed.



