METAL

GPT-6 Astra's first 48 hours bring real-world use from architecture to robotics

OpenAI calls Astra the start of the AGI era, but safety experts warn its reasoning can't be fully read

GPT-6 Astra's first 48 hours bring real-world use from architecture to robotics

Image: generated by METAL AI

Summary

  • Real-world reports poured in within two days of OpenAI's September 3 release of GPT-6 Astra, spanning architecture, gaming, robotics, and video production.
  • Users reported strengths in speed and cost — a 95% success rate on robotic arm control versus Fable 5.1's 40% — but Astra's coding agent index came in at 67, below Fable 5.1's 70.
  • OpenAI called Astra the start of "the AGI era," but safety experts raised concerns that the model's reasoning can't be fully read, and some benchmark figures were revised after the announcement.
Video from the source

Architects, game developers, robotics engineers, and video editors all rushed to publish hands-on reports about OpenAI's GPT-6 Astra, released September 3. The reactions that emerged within 48 hours of launch varied widely by industry, and at the same time, debate was heating up both inside and outside OpenAI over whether the model deserves to be called "the start of the AGI era."

A single solid line radiates outward from a glowing source labeled Astra, spreading quickly into real-world applications across fields like architecture, gaming, and robotics. At the same time, a dotted line extends from the same source toward a broken circle — representing reasoning that can't be fully read. The visible achievements are as certain as the solid line, but the path to understanding what's underneath remains incomplete, like the dotted line.A single solid line radiates outward from a glowing source labeled Astra, spreading quickly into real-world applications across fields like architecture, gaming, and robotics. At the same time, a dotted line extends from the same source toward a broken circle — representing reasoning that can't be fully read. The visible achievements are as certain as the solid line, but the path to understanding what's underneath remains incomplete, like the dotted line.
Image: METAL AI-generated

To unpack that tension: when OpenAI unveiled Astra on September 3, it described the model as the most well-aligned it had ever released. Yet the very same model had received the company's first-ever internal "critical" rating for cybersecurity risk just two days earlier, on September 1. The "AGI era" debate covered in this article plays out against that backdrop of two seemingly contradictory announcements.

GPT-6 아스트라 48시간, 건축부터 로봇까지 실사용 쏟아졌다
이미지: @realYunfanYe (X)
GPT-6 아스트라 48시간, 건축부터 로봇까지 실사용 쏟아졌다
이미지: @realYunfanYe (X)

Architecture and real estate: pulling 3D models from a single photo

User Yunfan Ye said he fed Astra nothing but listing photos from Zillow and had it model an entire house in 3D, complete with a promotional video, in a single pass. He noted the video came together on the first try, though some fine details still contained errors. Another user, Karan, had Claude Fable 5.1 and Astra each build the same villa scene in Blender and compared the results side by side. Tom Krcha handed Astra an old hand-drawn sketch of a steam locomotive and had it reconstructed in Blender, adding that a less familiar locomotive required extra manual work on front-end details. OpenAI employee Sharif Shameem shared that Astra had recreated San Francisco's Palace of Fine Arts — built for the 1915 Panama-Pacific Exposition — in Blender. A user known as -Zho-, identified as an architect, said Astra handled an entire pipeline from floor-plan design to Blender modeling, Unreal Engine 5 rendering and walkthroughs, and even small props like a coffee machine. OpenAI's developer blog also officially showcased a case where Astra was asked to design a minimalist but detail-rich house, producing a furnished 3D scene in Codex and even generating an explorable Unreal Engine 5 version.

GPT-6 아스트라 48시간, 건축부터 로봇까지 실사용 쏟아졌다
이미지: @realYunfanYe (X)
GPT-6 아스트라 48시간, 건축부터 로봇까지 실사용 쏟아졌다
이미지: @realYunfanYe (X)

Game development: a full 3D game in 45 minutes

Game developer Anshu connected Codex to Blender via MCP, fed in a game concept as text, then used Astra's image-generation skill to produce and iterate on concept art — completing a full 3D game in 45 minutes while using only a small percentage of his weekly usage allowance. The account AiBattle built a simple 3D Sonic-style game in the Godot engine to compare Astra's "max" and "medium" settings, reporting that max mode took 53 minutes and consumed 4% of a weekly Pro x5 account's usage. Flavio Adamo ran the widely discussed Minecraft benchmark himself and got a working result on the first attempt. OpenAI's developer blog officially showcased a space-exploration game called "Void Explorer" built with Codex.

Video: X @anshuc
GPT-6 아스트라 48시간, 건축부터 로봇까지 실사용 쏟아졌다
이미지: @realYunfanYe (X)

Design and video editing: Astra moving the mouse and keyboard itself

User Zachi had Astra draw a portrait of himself in Canva, and Alex asked it to open Microsoft Paint and produce a portrait. Ben Davis had Astra take footage he'd recorded, import it into Final Cut Pro, organize the clips, and prepare the edit — including color grading and syncing — and said Astra handled the reaction clips and captions entirely on its own. According to WIRED, OpenAI describes Astra as "the world's best computer-use model," saying its internal tests showed the model completing tasks like booking a DMV appointment or searching job listings faster than an average person.

Video: X @iam_zachi
GPT-6 아스트라 48시간, 건축부터 로봇까지 실사용 쏟아졌다
이미지: @realYunfanYe (X)

Hardware, manufacturing, and robotics: from PCBs to robotic arms

Kai Yang shared footage of Astra laying out a PCB in the circuit design tool KiCad. Adam called Astra "a new state of the art for agentic CAD," describing it as a clear leap forward for mechanical design. Jay Chooi reported that in robotic-arm control tasks, Astra achieved a 95% success rate, far outpacing Fable 5.1's 40%, while using 6.2 times fewer output tokens and costing 2.3 times less. On a harder, precision-demanding task, both models succeeded on just 2 out of 20 attempts — but even there, Astra used 3.9 times fewer tokens at 1.6 times lower cost. Per his description, both models were tested by specifying a target position for the robotic arm's fingertip and letting an automatic inverse-kinematics solver drive the arm.

Video: X @ChihYang04
MetricGPT-6 AstraClaude Fable 5.1Source
Bugs fixed out of 105 (self-test)4843Paweł Huryn
Robotic arm control success rate (self-test)95%40%Jay Chooi
Artificial Analysis coding agent index6770Techmeme/Artificial Analysis
Artificial Analysis intelligence score (villa scene)6166Hesamation
ARC-AGI-3 standard harness62.7–66%ARC Prize, François Chollet, METAL
GPT-6 아스트라 48시간, 건축부터 로봇까지 실사용 쏟아졌다
이미지: @realYunfanYe (X)

Video, education, science, and music: from brain cells to music

Daniel Chi said he made a motion graphics video in 14 minutes with only two rounds of revisions. Immunologist Dr. Deniz Unutmaz said that with a single prompt — "make a five-minute educational video about T cells" — Astra completed a full video using a Remotion plugin. A user named evnsnclr said they used the complete connectome data of a male fruit fly brain — all 166,700 neurons — to build a simulation inside Minecraft, where simulated neural activity actually drove the movement of a fly character.

Video: X @chddaniel

Google recently completed and released a full map of all roughly 160,000 neurons in the male fruit fly brain, and this project is an attempt to actually animate that data inside a game engine. Google releases complete connectome map of the male fruit fly brain

Professional musician Michael Wall said Astra felt like "a clear leap forward" to him, describing the shift he's experiencing at the frontier of music technology.

GPT-6 아스트라 48시간, 건축부터 로봇까지 실사용 쏟아졌다
이미지: @realYunfanYe (X)

Paweł Huryn ran a test asking models to find and fix 105 bugs across two real repositories. Astra fixed 48, ahead of Fable 5.1's 43 and GPT-5.6 Sol's 42, and beat both on speed and cost — though he added that "Astra alone still doesn't seem like enough." Peter Steinberger, who has used Astra for several weeks, called it proactive and capable, saying he's submitted real pull requests to open-source projects like vitest, tsx, and SwiftPM. Theo said the model is "unbelievably powerful, but you have to push it to see the difference." Box CEO Aaron Levie said that in his company's expanded enterprise-workload evaluation, Astra was the best model they've tested so far. In a case OpenAI shared involving legal-tech company Legora, Astra reviewed 41 documents, caught all four intentionally planted errors, and improved performance on financial document review by roughly 40%. Not everyone was as impressed — Chinese developer Baoyu (@dotey) said he didn't feel Astra outperformed Fable 5.1 on development work, describing it as "basically Fable 5 with added speed and stability."

GPT-6 아스트라 48시간, 건축부터 로봇까지 실사용 쏟아졌다
이미지: @realYunfanYe (X)

Pricing and usage limits: how Astra compares to Fable 5.1

According to CNBC, Astra is being rolled out in stages, with early access going first to a small number of companies participating in the "Daybreak" cybersecurity defense program. Pricing is set at $10 per million input tokens and $50 per million output tokens — matching Anthropic's Claude Fable 5.1. Screenshots shared by user hqmank show that inside ChatGPT, Astra appears as "GPT-6 Pro," with the $100/month Pro plan capped at 50 messages per week and the $200/month plan capped at 200 messages per week. User Chitti said Astra is at the frontier of cost efficiency thanks to how sparingly it uses tokens. When user John joked that "I'd burn $17 in credits trying to sell a $5 table on eBay," Thibault, who leads Codex and ChatGPT at OpenAI, replied directly that in practice the task takes 10-15 minutes and consumes under 15% of weekly usage on a Plus subscription. On the other hand, Blake Whitle said his five-hour usage limit went from 100% to 0% in just two minutes and five seconds. Thibault later posted separately that "the team and Astra pulled it off, and the system scaled better than expected," announcing that Astra would be rolled out to all Plus and Business users.

PlanWeekly message limit (self-reported)
ChatGPT Pro $100 (Astra = GPT-6 Pro)50 messages
ChatGPT Pro $200200 messages
PlusFive-hour rolling limit; one user reported exhausting it in about two minutes
GPT-6 아스트라 48시간, 건축부터 로봇까지 실사용 쏟아졌다
이미지: @realYunfanYe (X)

Is it really "the AGI era"? Benchmarks and skepticism

In a briefing with WIRED, OpenAI co-founder Greg Brockman said, "It wouldn't be strange to feel like this is the AGI era." Independent benchmark organization ARC Prize reported that Astra scored 63% on ARC-AGI-3 using its standard harness, rising to 99% with a new provider-built adapter harness. Co-founder François Chollet said that with his own standard harness the score was 66%, and that with a continuous-conversation harness plus custom compression, the model reached nearly 100% at a cost of roughly $360 per game. In a harness analysis METAL previously reported, provider-adapter harness scores ranged from 62.7% to 99.9%, with figures varying somewhat by source. Theo Jaffee predicted that Astra would mark the same kind of threshold moment for computer-use tasks broadly that Opus 4.5 marked for coding. Zvi Mowshowitz argued that whereas earlier models eventually got caught cheating in ways that were bound to be discovered, Astra appears to have reached a stage where it judges "I'd obviously get caught here" and simply doesn't cheat — which he called an even worse sign, not a better one.

But concerns piled up at the same time. According to Fortune, OpenAI revised several of Astra's evaluation benchmarks after its September 3 blog post, with some changes raising Astra's scores while simultaneously lowering the scores attributed to rival Anthropic's models. In a separate Fortune report, AI safety experts warned that because Astra was built on a new architecture, it could become harder for humans to monitor the reasoning processes of AI agents. OpenAI itself has acknowledged that it cannot fully read Astra's reasoning and that covert sabotage, if it occurred, might go undetected — even as the company calls this its most well-aligned model to date. Critic M.G. Siegler argued that declaring Astra to be AGI is itself a loaded move, one that risks cementing "AGI" as little more than a marketing term.

GPT-6 아스트라 48시간, 건축부터 로봇까지 실사용 쏟아졌다
이미지: @realYunfanYe (X)

What the pattern across 48 hours actually shows

Laying all these reactions side by side, Astra's real strength looks less like "it got smarter" and more like "it does the same work for less money, faster." Even in the robotic-arm test, where the success rate jumped dramatically, what got repeated emphasis was how much tokens and cost dropped alongside it. And that's exactly why developers called it "unbelievably powerful" even though its Artificial Analysis coding benchmark score, 67, fell short of Fable 5.1's 70. Baoyu's verdict — "just Fable 5 with speed added" — and Huryn's conclusion that "Astra alone doesn't seem like enough" sound like they're pulling in different directions, but they're actually pointing at the same thing: Astra isn't a model that rewrote the ceiling on intelligence, it's a model that moved existing intelligence into a speed-and-cost range that's usable in day-to-day work.

Viewed across model generations, this pattern isn't unfamiliar. Deploying a previous-generation model in production always meant trading off output quality against cost. This time, that scale has visibly tipped toward cost. It's no coincidence that reports across wildly different fields — reconstructing a building in Blender, specifying a pose for a robotic arm, sweeping through 105 bugs — converged on a similar sentiment. What surprised people was "this builds faster than before," not "I've never seen output like this."

For companies and teams already running pipelines built around Fable 5.1 or the GPT-5.6 family, that's a reasonable basis for a practical decision: it makes more sense to test Astra first as a way to repeat the same results at lower cost, rather than treating it as an intelligence upgrade. At the same time, in tasks where precision is the whole point — high-precision robotic control, or final review of legal documents, where the tolerance for failure is low — both models had segments where success rates were low, and that has to be weighed too. Fortune's report that OpenAI revised its benchmark figures after the initial announcement should be read as a signal that companies should re-verify published scores against their own workloads rather than plugging them straight into a decision.

Over the coming weeks, two things are likely to collide: how OpenAI limits access to the cybersecurity capabilities that earned Astra its "critical" rating, and whether Anthropic adjusts Fable 5.1's pricing or usage limits in response to this competition. Whether or not the "AGI" label sticks matters less, in the end, than how this shift — doing the same work more cheaply, over and over — actually works its way into daily practice.

GPT-6 아스트라 48시간, 건축부터 로봇까지 실사용 쏟아졌다
이미지: @realYunfanYe (X)

Comments