METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Anthropic Releases Claude Sonnet 5.5

Anthropic released Sonnet 5.5, the second model in the Claude 5.5 family. It generates output more than 30% faster than Sonnet 5, costs up to 30% less per task, and comes close to Opus 5.5 on several benchmarks.

Anthropic Releases Claude Sonnet 5.5

Image: @claudeai (X) (video still)

Summary

  • On September 28, Anthropic released Claude Sonnet 5.5, keeping token prices the same as Sonnet 5 while cutting cost per task by up to 30%.
  • It scored 70.6% on Terminal-Bench 4.0 and 1844 on GDPval-AA, close to Opus 5.5, and at Low or Medium effort it beat Sonnet 5's best score for about a tenth of the cost.
  • It is the first Sonnet model with cyber safeguards and classifiers that block reasoning extraction, and it is available on all platforms as claude-sonnet-5-5.
Introducing Claude Sonnet 5.5

On September 28, Anthropic released Claude Sonnet 5.5, the second model in the Claude 5.5 family. In its announcement, the company said Sonnet 5.5 generates output more than 30% faster than Claude Sonnet 5 and costs up to 30% less per task for most work. Token prices stay exactly the same as Sonnet 5; the savings come from finishing the same work with far fewer tokens. Claude Haiku 5.5, built for high-volume and cost-sensitive use, will join the family in the coming weeks.

The division of labor between the two models is clear. Anthropic presented Opus 5.5 as the model for complex work that requires careful judgment, and Sonnet 5.5 as the model for well-scoped everyday tasks, bug fixes, and documents, slides and spreadsheets. The company added that it has a sharp eye for design. METAL has reported on the September 22 release of Opus 5.5, the first model in the family, and the mid-tier model followed six days later.

The benchmark table shows the gap with the higher tier narrowing sharply. On Terminal-Bench 4.0, which measures multi-step tasks in a command-line interface, Sonnet 5.5 scored 70.6%, far above Sonnet 5's 10.3% and even higher than the 66.4% Opus 5.5 posted at its highest setting. On GDPval-AA, which tests real-world work across 44 occupations, it scored 1844, two points below Opus 5.5's 1846 and ahead of Sonnet 5's 1449 and GPT-6 Sol's 1487. On the computer-use test OSWorld 2.1 it reached 80.1%, close to Opus 5.5's 81.8%, and on the chart-reading test Chartography it jumped from Sonnet 5's 15.6% to 61.6%. On CursorBench 4.0, built from tasks in real Cursor coding sessions, it scored 55.5%, about two points behind Opus 5.5's 57.8%. It is also the first Sonnet model to beat Pokémon Red working only from screenshots.

What Anthropic emphasized more was cost per point. In the accuracy-versus-cost curves the company published, Sonnet 5.5 at Low or Medium effort beat Sonnet 5's best score on several benchmarks for about a tenth of the cost per task. On FrontierCode, which measures whether a code change could be merged, it scored 10 points higher than Sonnet 5 at the same High setting for about one fifteenth of the cost per task. The default effort is Medium in the Claude apps and Claude Code, and High on the Claude Platform. Turning effort all the way up did not always help. On FrontierCode, the Max score of 46.2% was lower than the 52.1% at Xhigh, one level below, and the company explained that at Max the model sometimes split code review across many subagents, which led to timeouts or edits beyond the task's scope. Anthropic wrote that Opus 5.5 remains clearly stronger at open-ended work requiring sustained judgment.

앤스로픽이 공개한 벤치마크 표. Sonnet 5.5, Sonnet 5, Opus 5.5, GPT-6 Sol의 Terminal-Bench 4.0, FrontierCode, CursorBench, GDPval-AA, AA-Briefcase, Humanity's Last Exam, OSWorld 2.1, Chartography 점수를 비교한다

Early customer numbers point to fewer steps and fewer tokens. Gabriel Grinberg, AI Engineering Lead at Base44, said, "Across 118 real app builds, Claude Sonnet 5.5 produced apps that scored level with Opus 5," adding, "It got there in 3.6 iterations per build on average, where Opus 5 took 7.7." Balyasny Asset Management said that across 2,441 finance tasks Sonnet 5.5 used about 121,000 tokens per answer, far fewer than Sonnet 5's 497,000. Fabian Hedin, co-founder and CTO of Lovable, said its coding evaluations showed a third fewer tool calls and roughly half the shell runs.

Terminal-Bench 4.0 정확도와 과제당 비용 곡선. Sonnet 5.5는 Medium 설정에서 Sonnet 5의 최고점을 넘고, Max 설정에서 70% 선에 닿는다

Results from workplace tools point the same way. Curtis Allen, Principal Engineer at Slack, said, "Without changing any of our prompts, Claude Sonnet 5.5 did better than Sonnet 5 on almost all of our offline Slackbot evals," with fewer steps and about 14% fewer output tokens. Zendesk said that across hundreds of real support cases the model made fewer wrong decisions and processed tickets 20% faster, and Box said it was more accurate than the previous model, 2.4 times faster, and used 12% fewer total tokens. In an internal Anthropic test, the model was given a public company's quarterly earnings materials, call transcripts and a slide template and asked for a 10-slide operating review, and two experts judged its first draft ready to send as is.

The safeguards moved up a level. Anthropic said Sonnet 5.5's cybersecurity capabilities are comparable to Opus 5's, making it the first Sonnet model to launch with the cyber safeguards and fallbacks used for its most capable models. Routine bug finding and fixing still work, but higher-risk cybersecurity requests visibly fall back to Sonnet 5. To counter distillation attacks, in which thousands of fake accounts are used to extract a model's capabilities at scale, it is also the first Sonnet model with classifiers that block reasoning extraction, and Anthropic expanded preserved thinking so that the model's thinking cannot be decoupled from the account that created it. The biology safeguards are the same as Sonnet 5's. On an automated behavioral audit of roughly 1,850 scenarios, Sonnet 5.5 matched or improved on Sonnet 5 on most alignment measures, and in containment evaluations it was the least likely of any Anthropic model to probe the limits of its containers. According to reports, the model also includes invisible text watermarking to comply with the EU AI Act.

The model is available on all platforms, including AWS, Google Cloud and Microsoft Azure, and developers can call it as claude-sonnet-5-5. Like Opus 5.5 and Sonnet 5, it can be used with zero data retention. Developers who ran Sonnet with thinking turned off need to switch to the new between_tools setting before migrating. The announcement post METAL reviewed on Claude's official account carried a 13-second video and had passed 2.6 million views and 31,000 likes at the time of checking. The video includes a scene of a flock of birds filling the sky, and the announcement used a prompt asking for a murmuration of 400 starlings in one HTML file as its speed-comparison example.

The axis of price competition is also shifting. According to reports, the standard API price of GPT-6 Sol, which OpenAI released last week, is the same as Sonnet 5.5's, and Google is selling Gemini 3.8 Flash at a lower introductory price through the end of the year. Anthropic launched Sonnet 5 in June at introductory pricing and made those rates permanent in August, which METAL reported as Sonnet 5's introductory pricing becoming permanent. This time it left prices untouched and lowered costs by reducing the tokens and tool calls needed to finish a task.

This release changes how companies choose models. As the mid-tier's scores climb to just below the top tier, there is now a case for keeping much of the work previously sent up to Opus on Sonnet. The unit of comparison is also moving from the price of a single token to the total needed to finish a single job, and Anthropic has put that calculation forward itself, alongside its benchmarks.

Comments