
Image: @ClaudeDevs (X) (video still)
Summary
- Anthropic published a Claude Sonnet 5.5 developer guide on September 28, listing five breaking changes and one change to the response shape for code moving from Sonnet 5.
- Disabling thinking, forcing a tool call and the old computer-use declaration now return 400 errors; the guide recommends the between_tools setting and auto tool choice with strict tools.
- Sonnet 5.5 is 30% faster than Sonnet 5 and costs up to 30% less per task, and Claude Code points its sonnet alias at the new model from v2.1.284.
Anthropic published a developer guide for Claude Sonnet 5.5 on its developer blog on September 28. Its bottom line is that swapping the model name does not finish a migration. Code coming from Sonnet 5 has to work through five breaking changes and one change to the response shape, and many of them send the request straight back with a 400 error. A post announcing the guide from Anthropic's developer account, ClaudeDevs, on X the same day passed one million views.
Addy Osmani, who wrote the guide, said Sonnet 5.5 is 30% faster than Sonnet 5 and that, while the per-token price is unchanged, it finishes the same work with fewer tokens, cutting costs by up to 30% for most work. Metal previously reported Anthropic's release of Claude Sonnet 5.5, and this guide is the manual for people putting that model into real products. According to the announcement, Sonnet 5.5 scored 70.6% on the agentic coding evaluation Terminal-Bench 4.0, far above Sonnet 5's 10.3%.
The guide starts by splitting the work. It assigns well-scoped everyday coding such as fixing bugs, iterating quickly on features and checking against requirements, plus repeatable agent tasks like investigation, review and drafting, to Sonnet 5.5, and long-horizon agentic coding and knowledge work that needs careful judgment to Opus 5.5. Osmani wrote that Sonnet 5.5 fits best when a task has a clear spec and a way to check the result. Claude Haiku 5.5, aimed at high-volume, low-latency work, is due to join the family within weeks.
The first trap is thinking. Sonnet 5.5 answers with thinking on by default, so a response can begin with a thinking block, and code that reads the text of the first block directly breaks. The thinking disabled setting now returns a 400 error. The new between_tools setting instead limits thinking to the gaps between tool calls, but it only works at the low, medium and high effort levels and returns a 400 error again at xhigh or max.

Forced tool calls are closed off too. Setting tool_choice to any or to a specific tool returns a 400 error, even on the token counting endpoint, and the guide recommends switching to auto and marking the tool strict so its input follows the schema. Computer use works only through the new computer_toolset_20260801 toolset, and the old declaration becomes an error, although Amazon Bedrock still accepts the old form. With the advisor tool, a Sonnet 5.5 executor rejects Opus 4.8, Opus 4.7 and Sonnet 5 as advisors, and advice from accepted advisors comes back encrypted, so code cannot read it.
Thinking blocks are tied to the model and the conversation. Sonnet 5.5 reads Sonnet 5's thinking blocks, so switching models mid-conversation keeps the reasoning, but no other model can read Sonnet 5.5's blocks. The guide therefore asks developers to keep conversations append-only. Those who would rather not migrate by hand can hand the job to the /claude-api command in Claude Code, where the bundled Claude API skill swaps the model ID and the breaking parameters across the whole codebase at once.
Effort has to be measured again. The guide said effort levels were recalibrated, so a Sonnet 5 setting no longer produces the same amount of thinking. The default is high on the Claude API and medium in Claude Code. The guide recommends starting agentic coding at medium and raising only harder tasks to high, and starting latency-sensitive work such as chat at medium or low. It adds that asking for less thinking in the system prompt does not reliably reduce it, so the effort level should be lowered instead. Workarounds bolted onto Sonnet 5, such as telling it not to be lazy, should be removed and evals re-run first.
Costs were reworked as well. The minimum cacheable prompt drops from 1,024 tokens to 512, so shorter system prompts and tool definitions now qualify for caching. A cache read costs a tenth of the input price, but changing the top-level effort between requests invalidates the cache. Images use a high-resolution tier of up to 2576 pixels on the long edge, so a single 2000×1500 image uses about 2.5 times as many tokens as on Sonnet 4.6, and US-only inference costs 1.1 times the standard price. The context window is one million tokens with no beta header, maximum output is 128,000 tokens, and the knowledge cutoff is June 2026.
Outside developers focused on speed and token savings. Daniel Vogel, chief operating officer of Epic Games, said "the new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks." David Loker, vice president of AI at code review company CodeRabbit, said "we plan to move simple and moderate reviews over now, and more in the coming weeks."

Refusals are handled differently. A declined request comes back with HTTP 200 and a refusal reason in one of five categories: cyber, bio, frontier LLM, reasoning extraction and general harms. Anthropic said Sonnet 5.5 is the first Sonnet model with cybersecurity safeguards similar to those on its most capable models, and the server-side fallback retries only cyber and frontier LLM declines on Sonnet 5. To avoid reasoning extraction declines, the guide advises against asking the model to write out its reasoning in the response and recommends reading summarized thinking instead.
In the 40-second demo video in the guide, which Metal reviewed, the two models repaint the same photograph with Python code they wrote themselves. According to the on-screen text, no image model was used, and by the time the paintings were finished Sonnet 5 had called the brush engine it built 12,249 times and Sonnet 5.5 44,561 times, filling the city at dusk far more densely. The X post carries a reader-proposed context note, still being rated, saying the photographer who took the reference image stated it was used without credit or permission, while the guide itself credits that photographer's account for the reference image.
In Claude Code, from v2.1.284 the sonnet alias points to Sonnet 5.5 and runs at medium effort. Thinking cannot be turned off, there is no fast mode, and the default model remains Opus 5.5. The model is available on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry, where it is offered only on Global Standard deployments.
Seen through an engineer's eyes, this guide has a longer compatibility list than benchmark table. In exchange for a cheaper, faster model, the thinking default, forced tool calls, computer use and advisor pairings all changed at once, and Anthropic chose to ease that cost with the /claude-api command and a skill. The document's demand on developers is to check the list of 400 errors and re-measure effort before changing a single line of model ID.





Comments