METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Anthropic launches developer hub claude.dev

Anthropic has opened claude.dev, a developer hub that gathers engineering deep dives and Claude Code and API guides. Its flagship posts are the record of a sprint that made claude.ai about 3x faster in two weeks and a method for automating eval design.

Anthropic launches developer hub claude.dev

Image: @ClaudeDevs (X) (video still)

Summary

  • Anthropic officially opened claude.dev on September 30, a developer hub for people building with Claude that it says will carry engineering deep dives, Claude Code and API guides, and easter eggs.
  • According to the hub's flagship post, Anthropic merged more than 3,000 changes with Claude and no incidents during a two-week August sprint, making the core claude.ai experience about 3x faster.
  • The eval design post introduced the build-eval and hillclimb commands in the claude-api skill; in a customer support experiment, held-out accuracy rose from 78.6% to 90.5% while cost fell to about one fifth.
Claude.dev is our new home for developers building with Claude.

Anthropic officially opened claude.dev on September 30, a developer hub for people building with Claude. The company's developer account ClaudeDevs introduced claude.dev on X as "our new home for developers building with Claude," saying it will carry engineering deep dives, Claude Code and API guides, tips from the teams building Claude, and some fun easter eggs. The post passed 110,000 views in less than a day. Anthropic has effectively pulled its scattered technical writing onto a single address for developers.

The site's identity sits in one line on its front page. claude.dev describes itself as a place for sharing tips, tricks, and points of view from Anthropic's developers, and tags every post with its reading time in minutes. The 10 posts on the front page were published between May 20 and September 28, and most cover practical topics: a 21-minute breakdown of what a single task costs on Opus 5.5, the new rules of context engineering for Claude 5 generation models, and how the Claude Code team uses skills. Official YouTube videos sit further down, and the footer gathers links to Claude Academy, Discord, offline meetups, and the Reddit community.

The easter eggs are real. The claude.dev terminal page that METAL checked, as of a browser visit on October 1, opens with ASCII art spelling CLAUDE.DEV in dots, tells visitors to type /help to discover other commands, and loads the post list with /posts. The 30-second video attached to the X post shows gameplay from a pixel game called Clawd's Quest. In a setting where the sky forgot to wake up one night, an orange pixel character named Clawd collects five sparks, the sun rises, and a thank-you message appears on screen. Hiding a game in a developer site reads as a choice to look more like a community space than a documentation site.

The weightiest post on the hub is the record of making claude.ai three times faster in two weeks. According to the September 23 post by three engineers, Raymond Wang, Sam Attard, and Issac G., Anthropic made the core user experience of the claude.ai web and desktop app about 3x faster in a two-week August sprint. The team targeted four user journeys that account for 95% of activity. At the 75th percentile, time to a typeable page on a fresh load of claude.ai fell from 3.1 seconds to 0.55, starting a new Claude Code session fell from 0.8 seconds to 0.3, and loading a Claude Cowork cloud session fell from 2.6 seconds to 0.73. The team hit 12 of its 13 targets within three days.

The way the work was done stands out as much as the results. The team ran every thread in a single Slack channel and attached Claude to each one. Anthropic said it ran an internal research model roughly comparable to Opus 5.5 through Claude Tag in beta. Claude found bottlenecks, built benchmarks, put up changes, and watched deploys, while humans set goals, made tradeoffs, and approved every change. In that way the team merged more than 3,000 changes without a single customer-facing incident or rollback. On the busiest days more than 200 changes landed, and at one point more than 150 threads were running at once. The team introduced nearly 200 feature flags and cleaned up more than half of them within the sprint.

8월 13일과 8월 27일 실사용자 측정 75백분위를 비교한 도표. 앱 실행, 대화 시작, 대화 불러오기, 메시지 전송 네 여정 13개 측정값의 전후 시간과 개선 배수를 보여 주며 평균 3.1배 빨라졌다

The findings were specific. Claude found 6,900 React hooks and 900 store subscriptions in the composer's typing path re-rendering on every keystroke, and caught hidden reloads happening half a million times a day. The cause of code block highlighting freezing for about a second turned out to be the em dash. If a reply contained any character such as an em dash or a curly quote, V8 stored the whole string as UTF-16 and took the slower path, and Claude fixed it with a twenty-line change. The time long replies blocked the main thread fell from about 750 milliseconds to 200, and they held 120 frames per second from start to finish on a 120Hz MacBook.

The human role was tuning ambition rather than speed. According to the post, Claude by default was careful about scope and padded its estimates. When threads that had hit their targets slowed down, Attard left the same message in thread after thread: "Let's keep driving this down, the targets are not the stopping point. What's next? Be ambitious." Shelley Vohr, an engineer in the channel, called the model "a numbers demon." When the results were shared internally, Issac G. said, "You could not have convinced me this was possible even six months ago."

The other flagship post covers eval design. The September 28 post by Lance Martin said the claude-api skill now includes two commands, /claude-api build-eval and /claude-api hillclimb. The first interviews the user and builds an evaluation set inside the codebase, and the second raises the score by applying one change at a time. Cases are split into a training set and a test set that is never shown, and if the training score rises while the test score stays flat, the change is treated as overfitting and reverted.

The numbers in the example lean toward cost. In an experiment on an internal customer support benchmark of 44 tickets, with 30 used for the search and 14 held out, the starting point was Opus 4.8 at default effort with 74.4% decision accuracy and 4.6 cents per ticket. After the prompt was cleaned up, switching to Opus 5.5 on low effort produced 87.8% at 1.9 cents, and Sonnet 5 on low effort delivered 88.9% at about 1 cent. According to Anthropic, Opus 5.5 input and output tokens cost 20% less than on Opus 4.8, and cache reads cost 60% less. On the 14 tickets the search never saw, the final configuration scored 90.5%, higher than the original setup's 78.6%, at about one fifth of the cost. The claude-api skill's own evaluation score also rose from 66% to about 88%.

Seen from the seat of someone running an engineering organization, what this hub sells is not a feature but a way of working. Both posts reach the same conclusion. If something can be counted, Claude can improve it, so the human job is to add more things to measure and to lay down rollback mechanisms first. METAL previously reported on Anthropic's launch of Claude Opus 5.5, and the posts on claude.dev serve as an operating manual showing the procedures real teams use to put that model to work.

The company itself listed the remaining work. The 3x speed post said the 95th percentile, other journeys, and very long conversations still have room to improve, and promised a separate post on contributions made upstream during the sprint to projects such as Electron, Chromium, and Node.js. For claude.dev to become an address developers return to rather than a product announcement channel, the key will be how steadily follow-up records like these accumulate.

Comments