
Image: @OpenAIDevs (X) (video still)
Summary
- OpenAI unveiled the Decisions API in limited preview at DevDay 2026 on September 29; it takes a question, predefined answers and text or image context, and GPT-6 Luna returns one answer.
- In the official demo video, routing 10,000 customer requests took 150 milliseconds per request on the Decisions API versus 1.6 seconds on the Responses API.
- Per-call pricing, the limit on candidate answers and whether it can be tuned on a customer's own data have not been disclosed, and OpenAI said it will share details at broad rollout.
OpenAI unveiled the Decisions API, an API built only for making decisions, at DevDay 2026 on September 29. A developer supplies a question, a predefined list of answers and context in the form of text or images, and GPT-6 Luna returns one answer from that list. The OpenAI Developers account said in a post on X: "Give your app real-time decision-making with Decisions API, powered by GPT-6 Luna." The API is in a limited preview open to select customers, and OpenAI said a broad release is planned in the coming days.
In its DevDay recap, OpenAI said the API "enables real-time decision-making by focusing Luna's intelligence on a specific set of user-defined questions with finite pre-defined answers." The company named three uses: classifying content, routing requests and choosing an agent's next action. Instead of writing prose like a chat model, it returns only one of the options the developer defined, so an application can use the result directly as a branching condition without parsing the response. According to reports, each result comes with a confidence score.
The 28-second official demo video that Metal reviewed focuses on speed. It runs the same job, splitting a spreadsheet of 10,000 customer requests into three categories, Billing, Technical and Sales, on the existing Responses API and the Decisions API side by side. The processing time shown on screen is 1.6 seconds per request for the Responses API and 150 milliseconds per request for the Decisions API. In the footage, shown at 15 times real speed, the Responses API had processed 2,000 requests when the Decisions API finished all 10,000, and the video closes with the line "~10x faster decisions." According to reports, OpenAI gave reporters the same comparison of 150 milliseconds against 1.6 seconds. Measurement conditions such as region, input size and concurrency appear neither in the video nor in the announcement.
The engine, GPT-6 Luna, is the smallest model in the GPT-6 family and was released on September 22. Metal previously reported that OpenAI released GPT-6 Sol and Luna. OpenAI's developer documentation describes Luna as "our most efficient model for focused, high-volume tasks." According to the documentation, Luna has a 1,050,000-token context window, a 128,000-token maximum output and a knowledge cutoff of May 18, 2026. It accepts text and image input and produces text output. The same page lists fine-tuning for Luna as not supported.

OpenAI did not open the decision-model category. TypeSafe AI released Jev on September 15, a model that writes no prose and returns only typed answers with confidence values, and made it generally available on September 21. Metal previously reported that TypeSafe unveiled its decision model Jev. According to reports, the Decisions API is likely a response to Jev's rapid rise, and one analysis said OpenAI rushed the announcement to make DevDay. OpenAI has a related precedent: its Moderation API, which screens harmful content, has long returned per-category scores instead of prose. Those categories, however, are set by OpenAI, whereas in the Decisions API the developer writes them.
An OpenAI spokesperson said the company would share more "at broad rollout," according to reports. Per-call pricing, the number of candidate answers a single request can carry, and whether it can be tuned on a customer's own data have not been disclosed. Rate limits and regional availability have not been published either. When Metal checked OpenAI's developer documentation index on September 30, there was still no guide or API reference entry for the Decisions API. The official information on the API so far consists of the X post, the DevDay recap and the demo video.
Seen through the eyes of a tech lawyer, the weight of this API lies less in speed than in the closed list of answers. Because every permitted answer is defined in advance, an application can reject any answer it did not define, attach a different permission level to each branch and hand the case to a person when context is thin, according to analysis from the developer industry. The fact that every test case has a target category, making it possible to measure the cost of a wrong answer beforehand, was also cited as a strength. The same analysis noted that a constrained model can still pick the wrong valid answer, be swayed by misleading context or inherit the bias of the categories the developer wrote. It therefore recommended adding a holding option such as needs review or unknown and routing it to a separate process. The model makes the pick, but the developer writes the list of choices. Until pricing and a service-level agreement are published, the first place this API will settle is not payments or customer messaging but reversible forks such as sorting inquiries.
Sources
- OpenAI Developers (X) — Give your app real-time decision-making with Decisions API, powered by GPT-6 Luna. →
- The New Stack — OpenAI answers TypeSafe's Jev with a Decision API built on Luna →
- OrcaRouter — OpenAI's Decisions API: GPT-6 Luna Picks One Answer →
- AlphaSignal — OpenAI's Decisions API Gives Developers a Constrained GPT-6 Luna Router →
- OpenAI — GPT-6 Luna Model | OpenAI API →





Comments