METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

OpenAI launches Ultrafast speed tier

At DevDay 2026, OpenAI launched Ultrafast, which runs GPT-6 Astra at up to 300 tokens per second. API pricing is 6x the standard rate, and in ChatGPT Work and Codex it is available only on the new Pro 500 plan and Enterprise.

OpenAI launches Ultrafast speed tier

Image: @OpenAI (X) (video still)

Summary

  • OpenAI launched its premium speed tier Ultrafast at DevDay 2026 on September 29, delivering up to 300 tokens per second, up to 8x faster in Codex and up to 6x faster in the API.
  • The first model is GPT-6 Astra, API pricing is 6x the standard rate, and in ChatGPT Work and Codex it is available on the new Pro 500 plan, which offers 25x Plus usage, and on Enterprise.
  • It is slower than the 750 tokens per second of the August preview but attaches speed to the top model, and Ultrafast for GPT-6.1 Sol is coming soon.
This is Ultrafast. Our premium speed tier, Ultrafast offers up to 8x faster token generation (300 tokens per second) in Codex and up to 6x in the API.

OpenAI launched Ultrafast, a premium speed tier, at DevDay 2026, held in San Francisco on September 29. On its official X account, OpenAI said Ultrafast raises token generation speed by up to 8x in its coding tool Codex and up to 6x in the API, reaching up to 300 tokens per second. The first supported model is its top model, GPT-6 Astra, available from the same day in the API, ChatGPT Work and Codex. Ultrafast for GPT-6.1 Sol is coming soon.

The speed is not free. According to reports, using Ultrafast in the API costs 6x the standard rate for the same model. OpenAI's existing Fast mode costs 2x the standard rate and, for Astra, runs about 2.5x faster. Ultrafast sits another 3x above Fast mode in price. The OpenAI Developers account wrote that in Codex, Ultrafast "runs up to 8x faster than Astra Standard and 4x faster than Astra Fast," adding, "Bring your ideas to life as fast as you can type them."

Using Ultrafast in ChatGPT Work and Codex requires a new plan. On the same day OpenAI introduced Pro 500, bundling its highest usage allowance, 25 times that of ChatGPT Plus, with access to Ultrafast. Business customers use it through the Enterprise plan. According to reports, the existing Pro plan's allowance in ChatGPT Work and Codex drops from 20x Plus to 10x, and GPT-6 Pro messages in chat are halved from 200 to 100 per week. Existing subscribers keep their current limits for a transition period and receive a one-time credit.

The 45-second launch video Metal reviewed is a race in which the two tiers build the same 3D scene side by side. On both the Standard screen on the left and the Ultrafast screen on the right, two robot arms on a circular workbench begin assembling a rocket. The Ultrafast side attaches the body and fins in 6.4 seconds, moves to a launch pad scene and sends the rocket into the sky. The Standard side is still at an empty workbench at that moment and only finishes assembly at 30.3 seconds before moving to the launch pad. The first post carrying the video had drawn about 878,000 views as of September 30.

The terms for developers are laid out in OpenAI's API guide. Developers set service_tier to ultrafast in the request, and GPT-6 Astra is open to all API users but starts with low rate limits. Default limits are 500,000 tokens per minute for usage tiers 1 to 3, 1 million for tier 4 and 5 million for tier 5. OpenAI strongly recommended WebSockets, which keep a connection open, for agents that make frequent tool calls, because opening a new connection for each request lets network latency eat into the speed gains. Ultrafast supports only US data residency and global processing, not EU or other regional processing endpoints.

오픈AI Ultrafast 발표 영상 장면. 같은 로켓 3D 장면을 만드는 경주에서 오른쪽 Ultrafast는 6.4초에 조립을 끝내고 로켓을 띄웠고, 왼쪽 Standard는 29.5초에도 로봇 팔이 조립 중이다

The tier had a preview a month and a half earlier. Metal reported on August 13 that OpenAI previewed an ultrafast mode running GPT-5.6 Sol up to 14x faster. At the time, OpenAI reached up to 750 tokens per second on Cerebras hardware and opened it only to a limited group of customers including Jane Street, Podium, Basis and Rogo. John Crepezzi, who works on AI Assistants at Jane Street, said then that "the increase in speed brought by Cerebras is impressive," adding that "it enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them." The API guide still lists GPT-5.6 Sol as a preview model.

This launch stepped back on raw speed and stepped up on the model. At 300 tokens per second it is slower than the 750 of the August preview, but the speed is attached to OpenAI's top model. According to reports, output speeds measured by benchmarking firm Artificial Analysis are about 201 tokens per second for Google's Gemini 3.5 Flash, about 769 for the speed-focused model Mercury 2 and about 1,491 for Celeris-1. GPT-5.3-Codex-Spark, which OpenAI earlier ran at more than 1,000 tokens per second on Cerebras, was also a smaller model with weaker benchmark scores. According to reports, what sets Ultrafast apart is not raw speed but delivering that speed on a top-tier GPT-6-class model.

Ultrafast was one of more than 20 announcements at DevDay. Metal also covered GPT-6.1 Sol and the always-on dots agents from the same event. Placed side by side, the announcements reveal OpenAI's plan. GPT-6.1 Sol brings near-Astra performance down to one-fifth of Astra's standard price to make repeated work cheap, while Ultrafast does the opposite, charging 6x for the moments when a person is waiting at the screen. In its August announcement, OpenAI listed incident response, financial research, customer support and commerce as use cases and wrote that "when speed no longer requires giving up intelligence, AI can move into the most time-sensitive parts of a business."

From a journalist's point of view, what received a new price in this announcement is not the model but the wait. Even when the same GPT-6 Astra produces the same answer, an answer that takes 30 seconds and one that takes 6 seconds now have different prices. Individual developers must move up to Pro 500 to buy that difference, and existing Pro subscribers see their limits halved. As speed becomes a product, OpenAI's price list has gained a third axis, time, alongside intelligence and cost.

Comments