Anthropic has launched Claude Haiku 5.5, updating the smallest and cheapest tier in the Claude family. Haiku sits below Sonnet and Opus, and is aimed at high-volume, cost-sensitive tasks where speed matters.

The biggest change is pricing. According to Anthropic, requests under 100,000 tokens cost $0.10 per million input tokens and $0.50 per million output tokens, which the company says is 90% cheaper than Haiku 4.5. Longer requests above that threshold cost $0.50 and $2.50 per million tokens, or about half the price of the previous model.

Because roughly 90% of requests on the previous Haiku were under 100,000 tokens, Anthropic estimates that average running costs are about 75% lower. A 100,000-token context is roughly the length of a novella, covering many everyday workloads such as customer support, document summarization, classification, and simple queries.

The model is also described as a step up in capability. In Anthropic’s published benchmarks, Haiku 5.5 improved from 15.7% to 72.4% on OSWorld 2.1, a test of multi-step computer use, and from 0% to 39.2% on Terminal-Bench 4.0, which measures agentic coding in the command line. The company says Haiku 5.5 is its fastest Claude model, except when Opus is used with fast mode enabled.

Anthropic points to two main kinds of use cases. One is high-volume repetitive work such as summarization, classification, database queries, real-time customer support, and browser automation. The other is acting as a sub-agent: letting Opus 5.5 or Sonnet 5.5 handle the main job while Haiku 5.5 takes on smaller sub-tasks.

Haiku 5.5 is also the first Haiku model with adjustable effort levels, with five settings from low to high. Developers can use lower settings for simple tasks to save cost, and higher settings when they need stronger performance.

Claude Haiku 5.5 is now available through the Claude API, as well as on AWS, Google Cloud, and Microsoft Azure, under the model name claude-haiku-5-5.

Alongside the launch, Anthropic also cut Sonnet 5.5 cache read pricing in half, to $0.10 per million tokens. Caching stores reused content such as long system prompts or full codebases, which can make up a large share of cost in agentic workflows. The company estimates that this change makes most Sonnet 5.5 agent tasks about 20% cheaper.