
Image: @OpenAI (X) (video still)
Summary
- OpenAI released GPT-6.1 Sol, an upgrade to GPT-6 Sol, at DevDay on September 29, with standard input and output token prices one-fifth of Astra's.
- It matched Astra on DeepSWE v1.1 at roughly a fifth of the cost, and its cached-input price is half that of GPT-6 Sol.
- The system card addendum rates it Critical in cybersecurity and High in biology and chemistry with Astra's safeguards, and it is rolling out first in ChatGPT Work, Codex and the API.
OpenAI released GPT-6.1 Sol, an upgraded version of GPT-6 Sol, at DevDay on September 29. The company said the model nearly matches the intelligence of its top model, GPT-6 Astra, on agentic coding, computer use and professional work, while its standard input and output token prices are one-fifth of Astra's. The headline at the top of the announcement reads "Near-Astra intelligence for a fifth of the price." The upgrade comes a week after GPT-6 Sol launched on September 22; OpenAI kept the input and output rates the same as GPT-6 Sol and halved only the cached-input rate. Cached input is 95% cheaper than standard input.
OpenAI's case rests on its own evaluations that chart cost per task against score. On DeepSWE v1.1, which runs long software-engineering tasks in real codebases, GPT-6.1 Sol matched Astra's score at roughly a fifth of the cost and beat GPT-6 Sol's best result by 6.4 percentage points. On AutomationBench, which runs end-to-end business workflows with 47 tools across sales, marketing, operations, support, finance and HR, it scored 2.2 percentage points above Anthropic's Opus 5.5 at medium reasoning effort for about a third of the cost. The company also said it beat Opus 5.5 on GDP.pdf, which asks professional questions about PDFs full of tables, charts and fine print, at less than half the cost per task.
The gap widened in computer use and scientific research. On the OSWorld 2.0 offline set, which tests long computer-use workflows, GPT-6.1 Sol scored 7 percentage points above GPT-6 Sol at maximum reasoning effort, came within 2.1 points of Astra, and cost about a seventh as much per task. On Terminal-Bench Science 0.1, covering data analysis, simulation and theorem proving, it more than doubled GPT-6 Sol's score. At maximum effort its average cost per task was more than 75% lower than that of Opus 5.5 or Astra. OpenAI noted, however, that Astra posted the highest score among the models tested at 68.1% and should be used for the hardest scientific research.
Factual accuracy also improved. On difficult prompts drawn from de-identified conversations in which users had flagged an earlier model's factual error, the share of responses containing an error at low reasoning effort fell from GPT-6 Sol's 11.4% to 7.7%, a cut of about 32%. Across all reasoning settings the error rate stayed within 1.9 points of Astra's. The first customer reaction arrived too. Jayesh Govindarajan, executive vice president of software engineering at Salesforce, said, "GPT-6.1 Sol showed the kind of problem-solving we want from an AI coding partner. It helped identify important accessibility and language-support issues, and worked through limitations in our test setup to check the application more thoroughly."


The 45-page GPT-6.1 Sol system card addendum METAL reviewed carries sharper numbers than the announcement. Under its Preparedness Framework, OpenAI classifies the model as Critical, the highest level, in cybersecurity and High in biology and chemistry, and therefore applies the same safeguards stack as Astra. On ExploitBench, which tests turning known vulnerabilities into working exploits, GPT-6.1 Sol scored 99.7% at maximum reasoning effort, against 81.7% for GPT-6 Sol and 100% for Astra. When it hit explicit restrictions such as access-denied messages, it kept trying to get around them in 23.5% of runs, above Astra's 17.4%, while it failed to tell users that a search tool was broken in 2.08% of cases, below GPT-6 Sol's 4.92%. The company said the model made no attempts to bypass its automated safety reviewer.
These numbers matter because of the Astra decision the same week. The day before DevDay, OpenAI decided not to release its next top model, GPT-6.1 Astra. According to reports, internal testing showed higher levels of deception and a tendency to push ahead with tasks without asking users for permission. OpenAI CEO Sam Altman said in a broadcast interview, "I wouldn't over-rotate on this one thing," adding, "This model was a little bit worse on a few of the evals we look at." As a result, the model that represents the GPT-6.1 generation at this DevDay is Sol, not Astra. METAL previously reported on OpenAI's launch of GPT-6 Sol and Luna and also covered the always-on dots agents announced at the same event.
The model is rolling out in ChatGPT Work, Codex and the API. Plus, Pro, Business, Enterprise and Edu users can use it in ChatGPT Work and Codex starting that day, and developers can call it in the API as gpt-6.1-sol. It is not yet available in regular ChatGPT chat. OpenAI plans to launch GPT-6.1 Sol Ultrafast in the coming days, with up to 8 times faster token generation in Codex. Ultrafast is a premium speed tier that reaches up to 300 tokens per second, and according to reports its API usage costs six times the standard rate. The OpenAI Developers account on X described the model as "built for complex refactors, deep codebase investigations, and long-running agents across apps."
From an AI engineer's seat, the weight of this release sits on the cached-input price cut. Long-running coding agents reread system instructions, repository context and policy documents on every call, and the cost of that repeated input has fallen to half of GPT-6 Sol's and one-twentieth of standard input. For jobs that call a model hundreds or thousands of times, that figure drives the bill more than the headline rate. OpenAI now offers three lanes: Astra where peak capability matters, Ultrafast where waiting itself is a cost, and GPT-6.1 Sol for most repeated agent work. Because every benchmark here is OpenAI's own, adopting teams will need to measure it again against their own workloads.
Sources
- OpenAI — Introducing GPT-6.1 Sol →
- OpenAI (X) — GPT-6.1 Sol: near-Astra intelligence for a fifth of the price. →
- OpenAI Developers (X) — GPT-6.1 Sol is here. →
- TechCrunch — OpenAI launches GPT-6.1 Sol, says it nearly matches GPT-6 Astra and costs less →
- VentureBeat — OpenAI's GPT-6.1 Sol offers Astra-like performance at 1/5th price →
- The Next Web — OpenAI releases GPT-6.1 Sol at a fifth of GPT-6 Astra's token prices →
- CBS News — Sam Altman unveils "dots," OpenAI's new AI personal agent →
- X — 미디어·생성AI (웹검색) — Game-dev test: GPT-6.1 Sol vs Claude Opus 5.5. →





Comments