METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Anthropic releases Claude Haiku 5.5

Anthropic launched its small model Haiku 5.5 and cut prices for requests of up to 100,000 tokens by 90% from its predecessor. It scored 72.4% on OSWorld 2.1, ahead of GPT-6 Luna in the same price tier, and is the first Haiku model with an adjustable effort setting.

Anthropic releases Claude Haiku 5.5

Summary

  • Anthropic released Claude Haiku 5.5 on Oct. 7. Per-token prices for requests of up to 100,000 tokens are 90% lower than Haiku 4.5, and the company estimates average running costs fall by about 75%.
  • In the launch benchmark table it scored 72.4% on OSWorld 2.1, 39.2% on Terminal-Bench 4.0 and 1620 on GDPval-AA v2.1, ahead of GPT-6 Luna (48.9%, 16.4%, 1437).
  • The same day Anthropic halved the cache-read price of Sonnet 5.5 and began giving monthly API credits to Max and Team subscribers.
Introducing Claude Haiku 5.5

Anthropic released its small model Claude Haiku 5.5 on Oct. 7. The company introduced it as "the cheapest, fastest, and most capable small model we've ever released." Per-token prices for requests of up to 100,000 tokens are 90% lower than for its predecessor Haiku 4.5, and Anthropic calculates that average running costs fall by about 75%. The same day it halved the cache-read price of Sonnet 5.5 and began giving Max and Team subscribers a monthly API credit.

Haiku 5.5 is the third model of the 5.5 generation, following Opus 5.5 on Sept. 22 and Sonnet 5.5 on Sept. 28. Anthropic said it is aimed at high-volume, repetitive work such as summaries, compaction, database queries and classification. For coding work, the company recommends using it as a subagent that takes on smaller tasks handed down by Opus 5.5 or Sonnet 5.5. As its fastest model at standard speed, it also suits live customer support and browser use, the company said. Anthropic noted in a footnote that it still runs slower than Opus in Fast Mode.

Pricing splits by request length. Requests of up to 100,000 tokens cost $0.10 per million input tokens and $0.50 per million output tokens, 90% below Haiku 4.5's $1 and $5. Requests above 100,000 tokens cost $0.50 and $2.50, a cut of only 50%. According to Anthropic, about 90% of requests to Haiku 4.5 were 100,000 tokens or shorter. The company explained that the 75% average figure also accounts for a new tokenizer that uses slightly more tokens for the same work. According to reports, the lower tier's input, output and cache rates are identical to those of GPT-6 Luna, the low-cost model OpenAI launched last month.

The benchmark table takes direct aim at rivals in the same price range. In the launch table METAL reviewed, Haiku 5.5 scored 72.4% on the offline subset of OSWorld 2.1, which measures operating a real computer. Haiku 4.5 scored 15.7% and GPT-6 Luna 48.9%. On Terminal-Bench 4.0, which measures command-line agentic coding, it scored 39.2%, ahead of Haiku 4.5's 0.0% and GPT-6 Luna's 16.4%. Its GDPval-AA v2.1 knowledge-work score was 1620, above Haiku 4.5's 735 and GPT-6 Luna's 1437. On Humanity's Last Exam it scored 45.9% without tools and 57.4% with tools.

A gap to the larger models remains. Sonnet 5.5, listed for reference, scored 83.9% on OSWorld 2.1 and 70.6% on Terminal-Bench 4.0. Anthropic drew the line clearly: Sonnet 5.5 and Opus 5.5 remain better for complex agentic coding, while Haiku 5.5 fits narrowly scoped tasks. All of these figures were measured by the company itself. According to reports, the Terminal-Bench score in the 39% range comes from the maximum effort setting, while the default medium setting scores about 20%.

Haiku 5.5, Haiku 4.5, GPT-6 Luna, Sonnet 5.5의 지식 노동·컴퓨터 조작·추론·에이전트 코딩·시각 추론 벤치마크 비교표

That effort setting is the model's new dial. Haiku 5.5 is the first Haiku-class model with an adjustable effort level. Users choose on the same model whether to save cost or push intelligence. The launch's OSWorld 2.1 chart shows score and cost per attempt rising together across five levels from Low to Max, with every level sitting above the GPT-6 Luna curve at similar cost.

Companies that tested the model before launch pointed to speed first. Aaron Vinh, a staff software engineer at Asana, said that in evaluations for AI Teammates, task-completion latency fell by more than 30% and inference per agent turn ran up to 2.5 times faster than with the model it uses today, calling it "a noticeably snappier experience." HubSpot said Haiku 5.5 posted the best score it has seen on its CRM suite, 92.8% averaged over three runs. AlphaSense, whose document question feature makes about 8 million calls a week, said the model scored 0.84 across 400 queries, a statistically significant improvement over Haiku 4.5's 0.76. Box said it scored 11 points higher than Haiku 4.5 at about half the latency.

OSWorld 2.1 오프라인 부분집합에서 노력 단계별 점수와 시도당 비용을 비교한 도표. Haiku 5.5 곡선이 GPT-6 Luna보다 위에 있다

The subagent use case is illustrated by financial AI company Rogo. While a larger model builds a presentation, a Haiku 5.5 subagent goes into a 10-K filing and pulls the one segment revenue line the deck needs. "It's accurate enough that we'd trust it there, and fast and cheap enough that we can run it a lot," said Alex Wang of Rogo. Cognition said that pairing Opus 5.5 as the lead with Haiku 5.5 as the sidekick in Devin Fusion held a FrontierCode score of 66.2 while cutting cost and latency.

Safeguards were reworked as well. Anthropic said the model improved substantially over Haiku 4.5 on most alignment evaluations and showed a lower willingness to cooperate with misuse. Its cybersecurity safeguards are stricter than Haiku 4.5's but allow a wider range of defensive work than Sonnet 5.5's. Techniques more likely to be used by attackers, such as penetration testing, remain blocked. Organizations that need broader access can apply to the Life Sciences Verification Program and the Cyber Verification Program.

Prices across the rest of the lineup moved too. Sonnet 5.5's cache-read price fell from $0.20 to $0.10 per million tokens. Because cache reads make up a large share of token consumption, the company estimates the cost of most agentic work falls by about 20%. Starting this week, Max 5x subscribers receive $100 a month in API credits, Max 20x subscribers $200, and Team subscriptions up to $500 pooled across users. The Python and TypeScript SDKs gained beta support for computer use and browser use. METAL previously reported on Anthropic's Sonnet 5.5 release and OpenAI's GPT-6 Sol and Luna release.

Haiku 5.5 is available now on the Claude Platform as claude-haiku-5-5 and on AWS, Google Cloud and Microsoft Azure. The weight of this announcement lies less in a single benchmark table than in price and division of labor. In a setup where a large model decides and a small model runs repetitive work countless times, Anthropic has pulled the price of the most frequently called slot down to the level of its rival's low-cost model. The open question is how close real bills will come to the company's 75% figure once developers split their work by effort setting and request length.

Comments