Key Points
- Anthropic launched Claude Haiku 5.5, priced about 75% below Haiku 4.5 on average to run.
- The model is the first Haiku class with an adjustable effort setting across five levels.
- Cheaper small models make large-scale AI agent deployment viable on far smaller budgets.
The latest:
Running costs for Anthropic’s smallest model class fall by roughly 75% on average under Claude Haiku 5.5, announced October 7, 2026. The company positioned it as its fastest and most capable small model, aimed at high-volume work including summarization, classification, database queries and subagent tasks. It is available immediately through Amazon Web Services, Google Cloud, Microsoft Azure and the Claude Platform API.
Details:
- The pricing: According to Anthropic, input tokens cost $0.10 per million for prompts up to 100,000 tokens and $0.50 above that, against a flat $1.00 for Haiku 4.5. Output pricing is $0.50 and $2.50 respectively, compared with a flat $5.00 previously.
- Caching rates: Cache reads are priced at $0.01 per million tokens below the 100,000-token threshold and $0.05 above it, down from $0.10 flat. Cache writes fall to $0.125 and $0.625, against $1.25 under the previous model, according to the company’s published rate card.
- The savings split: Anthropic said the cost reduction is uneven by request size: roughly 90% lower for requests under 100,000 tokens and about 50% lower above that threshold. The company also noted Haiku 5.5 consumes slightly more tokens per task because of an updated tokenizer, which offsets part of the headline saving.
- The effort dial: Haiku 5.5 is the first model in its class with an adjustable effort setting spanning Low, Med, High, Xhigh and Max, allowing users to trade cost against intelligence per request rather than switching models.
- Benchmark jumps: On Anthropic’s own published comparisons against Haiku 4.5, the model scored 72.4% versus 15.7% on OSWorld 2.1 for computer use, and 57.4% versus 18.7% on Humanity’s Last Exam with tools enabled.
- Agentic coding: The widest reported gap is on Terminal-Bench 4.0, an agentic coding benchmark, where Anthropic recorded 39.2% for Haiku 5.5 against 0.0% for Haiku 4.5. On GDPval-AA v2.1 the company reported 1620 against 735.
- Customer testing: Asana staff software engineer Aaron Vinh reported latency reductions of more than 30% on task completions. HubSpot’s Ze’ev Klapow said the model averaged 92.8% across three runs on the company’s evaluation suite, its best result to date.
- Box results: Yashodha Bhavnani, Box vice president of AI products, said early testing put Haiku 5.5 “11 points higher” than Haiku 4.5 at roughly half the latency. Anthropic did not specify which internal benchmark produced that figure.
- Target workloads: Anthropic named live customer support and browser use as speed-sensitive applications alongside the cost-sensitive batch tasks, signalling the model is pitched at production deployments rather than experimentation.
Background:
Haiku is Anthropic’s smallest and cheapest model tier, sitting below the Sonnet and Opus lines. Small models carry the bulk of repetitive, high-volume enterprise traffic, which makes per-token pricing the decisive factor in whether agent deployments scale commercially.
Between the lines:
The two-tier pricing structure concentrates the discount in short prompts: 90% cheaper below 100,000 tokens versus 50% above. That design rewards the exact workload profile Anthropic named as the target — high-frequency, short-context calls such as classification and subagent tasks — while leaving long-context work comparatively expensive. The company’s own disclosure that the new tokenizer consumes more tokens per task means realized savings will land below the headline figure.
What’s next
Independent benchmark replications of the OSWorld and Terminal-Bench figures, since all performance numbers so far come from Anthropic or its customers. Also watch whether rival providers respond with matching cuts to their own small-model tiers.