METAL

Tag

벤치마크

AI

Walden Yan Appears in Video on Devin Testing With GPT-6 Astra

The OpenAI Developers account posted a one-minute video of Cognition co-founder Walden Yan on September 15. Inside Devin, GPT-6 Astra backs up the claim that code works with tests, cutting the amount of code people have to read themselves.

By 김현국

30

AI

Tencent Hunyuan Releases EvolveScaler Paper

The Tencent Hunyuan team released the EvolveScaler paper and project page on September 15, an information-evolution data framework. On data built by defining state in code first and then rendering it as natural language, the median avg@5 across 14 models fell to 11.3 on the longest tier.

By 김현국

00

AI

Specific Releases Real-SWE Enterprise Code Benchmark

Specific released Real-SWE on September 12, a benchmark that measures AI coding models on private production codebases licensed from real companies. Fable 5.1 led with a 38.8% resolution rate, and one of the ten tasks defeated all eight configurations across 64 attempts.

By 김현국

30

AI

Sakana AI Ships Two Models That Pick Other Models

Fugu Max routes each task to the leanest model that can solve it and undercuts rival output pricing by 40 to 60 percent. Fugu Ultra v2 posted its best scores with Fable 5.1 and GPT-6 Astra kept out of its pool.

By 김현국

20

AI

AI-designed proteins split prediction from lab results

Of 1,320 mini-protein binders Claude designed, 354 actually bound — hitting 14 of 15 targets. One target that stumped human competitors, who managed only a 3.7% hit rate, yielded 40% for Claude. But picking winners before synthesis is still the hard part.

By 김현국

170