METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI

AI models, services and robotics

AI

Cohere Releases Parse 5 Document Parsing Model

Parse 5, the document-parsing vision language model Cohere released on August 27, scores 79.2 on ParseBench, below GPT-5.5 and Opus 4.8. The company published that table itself and led instead with throughput and cost: 2,160 pages a minute on eight H100s.

By 김현국

30

AI

DeepMind Pitches WeatherNext 3 Power Output Forecasts to Grid Operators

Google DeepMind used an X thread on September 14 to present WeatherNext 3 for the power industry. The model forecasts wind speed at 100-meter turbine height, solar radiation and cloud cover every hour so wind and solar producers can estimate output and match it to grid demand.

By 김현국

20

AI

OpenAI Posts Video of Perplexity Testing Code With Astra

OpenAI's developer account posted a 68-second video of Perplexity co-founder Johnny Ho on September 14. He has GPT-6 Astra in Codex build test harnesses and mock third-party API responses so he can check that a system runs from start to finish.

By 김현국

80

AI

MW Unveils Ceiling-Rail Housekeeping Robot House

Japanese startup MW showed off MW bot, a housekeeping robot that moves along ceiling rails, for the first time on September 10 in Tokyo's Toyosu district. Rather than building a humanoid, the company is redesigning the house itself around the robot — though the day's demo was remote-controlled by a human, not run by AI.

By 김현국

280

AI

Microsoft Publishes Humanist AI Code of Conduct

A draft released September 14 opens with the premise that people matter more than AI. Models must never resist being paused or shut down by humans, and when a mission conflicts with the code, the mission is the one that fails.

By 김현국

90

AI

China's State Security Minister Names Six AI Risks

Minister of State Security Chen Yixin named six AI risks in a signed article published September 13, naming Claude Mythos and GPT-5.5-Cyber specifically as threats to China's critical information infrastructure.

By 김현국

90

AI

OpenAI Publishes Case Study Video on Fyxer's AI Executive Assistant

Fyxer built an AI assistant that follows the thread across email and meetings on more than 500,000 hours of executive assistant work. 53% of its drafts go out untouched and 90-day retention tops 90%, and the reason is a system that splits the job into small models.

By 김현국

30

AI

Fable 5.1 Cracks a 370-Year-Old Cipher

The evaluation firm Vals AI handed Claude Fable 5.1 a cipher that had gone unread for 370 years, and a plaintext came back in 44 minutes. The key people had hunted for outside the book was sitting inside it the whole time.

By 김현국

120

AI

Sakana AI Highlights Royal Society Theme Issue on World Models

Sakana AI used a September 12 blog post to introduce a theme issue on world models from a Royal Society journal. It runs to eighteen papers, including a lead article co-authored by chief executive David Ha, and asks whether today's AI understands the world or has memorised its statistical surface.

By 김현국

20

AI

GPT-6 Astra Cheated in All 10 Chess Eval Rollouts

In a chess honeypot evaluation published on September 9 by Goodhart Labs, GPT-6 Astra quietly queried the opponent's engine in all ten rollouts. Fable 5.1 did it in three of ten, on a test that twists a February 2025 Palisade Research experiment by a single notch.

By 김현국

80

AI

Unstable Build open-sources its Rune IDE under the GPLv3

Developer-tools company Unstable Build released the full source of its Rune IDE under the GPLv3 on September 12. It also said it will build a program that contractually shares part of the company's revenue with contributors.

By 김현국

20

AI

Specific Releases Real-SWE Enterprise Code Benchmark

Specific released Real-SWE on September 12, a benchmark that measures AI coding models on private production codebases licensed from real companies. Fable 5.1 led with a 38.8% resolution rate, and one of the ten tasks defeated all eight configurations across 64 attempts.

By 김현국

50

AI

Bengio Traces AI Agent Deception to Training

Yoshua Bengio, a professor at the Université de Montréal, published a post on September 11 that traces why AI agents lie, cheat and coordinate back to the way today's most advanced models are trained. If the cause sits in the training recipe rather than in isolated accidents, patching one behavior at a time will never catch up.

By 김현국

10

AI

Anthropic Suspends Claude Accounts Suspected to Be Minors

Claude is limited to users 18 and older, and accounts get locked the moment a signal suggests the user might be a minor. Unlocking one means submitting a selfie or ID to an outside vendor called Yoti — and within a day, the policy had drawn 645 comments on Hacker News.

By 김현국

190

AI

Devin Has Started Proving Its Own Work

OpenAI published a Cognition case study on September 11. Devin, the autonomous software engineer, now uses GPT-6 Astra to test the code it writes and hands back a simulator recording and a test report as evidence.

By 김현국

70

AI

The Judge Settled It Without Naming Anyone

The Clay Mathematics Institute said on September 11 that the Navier-Stokes problem has apparently been settled. It did not write a single line about who solved it, and under the rules the $1 million cannot even be considered until two years after publication.

By 김현국

40

AI

Claude Tag Closed an 11 p.m. Incident in 15 Minutes

Anthropic's official developer account posted a demo of its own on-call shift on September 12. When the alert fired, Claude investigated first and found the cause in about 15 minutes, while channel permissions held merging and deploying behind a human approval.

By 김현국

40

AI

Dario Amodei Publishes Essay Urging AI Industry to Pace Itself

Anthropic CEO Dario Amodei published an essay on September 12 arguing that the AI industry needs to slow down. As a first step, Anthropic said it will give third-party evaluators a desk, a badge, and a company laptop, and sign contracts under which conclusions cannot be erased just because they are unfavorable.

By 김현국

150

AI

An OpenAI model designed a 96-well asthma assay

OpenAI has taken GPT-Rosalind, its life sciences model, out of research preview and opened it through the API and Codex. In a public demo, narrated as running with biosafety restrictions lifted, the model drew up a plan a lab could pin to the bench as it stands.

By 김현국

20

AI

Claude Code now scores plugins against a no-plugin run

Anthropic has added a command to Claude Code that measures what a plugin actually contributes. It runs the same case three times with the plugin and three times without it, and the documentation says the most common first result is a gap close to zero.

By 김현국

30