3 min read

Claude Haiku 5.5 cuts prices, with a 100K-token threshold

Input pricing starts at $0.10 per million tokens. A short answer to a prompt over 100K tokens pays $2.50 per million output tokens. Cache rates also rise fivefold above the threshold.
Claude Haiku 5.5 pricing chart showing input rates rising from $0.10 to $0.50 and output rates from $0.50 to $2.50 per million tokens for prompts over 100K tokens.

Anthropic released Claude Haiku 5.5 on October 7, targeting repeated API work and coding subagents. The company estimates that it costs around 75% less to run on average than Haiku 4.5. Its lowest rates apply to prompts up to 100,000 tokens. Above that threshold, input and output rates are five times higher.

Agent requests can grow as conversation history and tool results accumulate. Haiku 5.5 accepts up to a million tokens of context, so an agent can continue working after its requests cross the 100K-token pricing threshold.

Prompt length determines the price tier

The model documentation lists the following Claude Platform API pricing, per million tokens:

Token type Prompts up to 100,000 tokens Prompts over 100,000 tokens
Input $0.10 $0.50
Output $0.50 $2.50
Cache read $0.01 $0.05
Cache write, 5-minute lifetime $0.125 $0.625
Cache write, 1-hour lifetime $0.20 $1.00

The Batch API offers a 50% discount on input and output tokens.

The output rate also depends on prompt length. A short answer to a long prompt falls into the higher tier. When evaluating an agent, inspect the context sent with each request, including the history and tool results it carries forward.

Anthropic's pricing footnote says 90% of Haiku 4.5 requests fell into the up-to-100K category. Its estimated 75% average saving already accounts for the new tokenizer. That historical request mix explains the emphasis on the starting rate; teams using Haiku for longer agent runs need to check their own distribution.

On Hacker News, minimaxir argued that agents would quickly exceed the cutoff, while classification and ordinary generation could benefit from the lower rates. In a separate r/ClaudeAI post, user ceramgcf reported reaching 100K–200K tokens even on tasks they considered simple. The post gives an early user's experience; it has no controlled comparison or independently checked logs.

Small, separate jobs are the clearest place to start testing the price cut. A worker that classifies one item or summarizes one document can receive only the material it needs. If a coding subagent carries file contents and test output forward, later requests grow. Measure those requests before estimating a whole run at the starting rate.

Reasoning effort changes the cost of a task

Haiku 5.5 defaults to adaptive thinking at medium reasoning effort. Set output_config: {"effort": "low"} to try less thinking on simple tasks, then check the answers against the same examples.

In Simon Willison's launch-day experiment, generating an SVG of a pelican riding a bicycle took seven seconds at low effort and five minutes nine seconds at max. The reported costs were about 0.094 cents and 3.38 cents respectively. He found a better bicycle frame at medium effort and above. Those measurements come from a single drawing prompt. Classification and extraction workloads need their own effort comparisons.

Token counts also change. Anthropic's migration guide says the new tokenizer counts the same text as approximately 30% more tokens than Haiku 4.5, with the increase depending on content. That applies to output as well as input. Recount existing prompts before applying the new rates to an old usage report, especially requests close to the threshold.

Haiku 4.5 migration includes request changes

The migration guide lists request changes alongside the new model ID, claude-haiku-5-5. Replace thinking: {"type": "enabled", "budget_tokens": N} with thinking: {"type": "adaptive"}; the old configuration returns a 400 error. Omit temperature, top_p, and top_k, and replace any final assistant message used as a prefill with a user turn.

If a response starts with a thinking block, code that treats the first block as answer text reads the wrong block. Select content blocks by their type. Thinking also counts toward max_tokens, so a small limit carried over from Haiku 4.5 can leave no room for the answer.

Haiku 5.5 does not support Priority Tier. Teams with a Haiku 4.5 Priority Tier commitment need to plan capacity separately.

Before migrating a workload, compare accuracy, time and cost across effort settings, and record how often requests exceed 100K tokens.