
Image: METAL
Summary
- On October 2, OpenAI published an official model guide explaining how to choose, instruct and operate the three models in the GPT-6 family for different workloads.
- The guide assigns Astra to the hardest reasoning, Sol to coding, research and computer use, and Luna to repeated work, and says cached input tokens cost up to 95% less.
- It features GPT-6 Astra cases from Harvey, Cognition, Hex and Invideo, with Invideo reporting roughly three times the success rate on color-grading tasks.
OpenAI on October 2 (local time) published an official model guide that sets out which kinds of work each of the three GPT-6 family models is suited for and how to use them. The guide covers, in turn, how to pick the model, reasoning effort and speed that fit a task, how to revise prompts, skills and repository instructions, how to steer long-running work that spans hours or days, and what to check before putting a workflow into production. OpenAI introduced GPT-6 as its most advanced suite of models yet, and framed the guide as covering everything from turning an idea into a working prototype to multi-step workflows that cross code repositories, databases and external APIs.
Model choice falls into three lanes. According to the guide, GPT-6 Astra is for the hardest reasoning work where maximum intelligence is needed, while GPT-6.1 Sol fits complex coding, research and computer use. GPT-6 Luna is meant for focused, everyday repeated work at scale with a clear goal, such as extracting invoice fields, classifying requests or producing structured summaries. OpenAI recommends treating model choice and reasoning level as a trade-off between intelligence and price, and comparing each model's pricing for the task at hand.
Reasoning effort comes in four tiers. Low is for routine tasks such as extracting facts or making small edits, medium for work that requires judgment such as planning a feature or comparing options, and high for difficult debugging, deeper analysis or careful review. The guide says extra high and max should be tested only when high falls short, and kept only if the improvement justifies the added time and cost. In the API, reasoning effort can be changed mid-conversation without breaking the cache. There are also two speed options. Fast mode, for places where response time matters such as chat apps or coding tools, delivers faster and more consistent responses at a higher per-token cost than standard processing, while Ultrafast speeds up token generation itself independently of reasoning effort and is available only for GPT-6 Astra.

The guide names prompt caching as the key lever for cutting operating costs. According to OpenAI, cached input tokens cost up to 95% less than uncached input tokens, depending on the model. Getting that benefit means putting stable instructions and reference material first, placing changing task details after them, and keeping tool definitions consistent. The guide says cost estimates for a complete workflow should include cache writes and long-context rates, and recommends compaction for longer conversations, which reduces context size while preserving the state needed to continue. Before deployment, it advises running representative tasks to measure task success, latency and cost per successful task.
Its advice on prompts can be summed up as less, but clearer. "Models have gotten much better at understanding nuance and ambiguity, so overly specific guidance can now hinder results where it previously helped," said Eric Provencher, who works on developer experience at OpenAI. The guide says to start with the result you want, who it is for, the relevant context and constraints, and what counts as done. Skill descriptions should be short and explicit about when each skill runs, supporting details should load only when needed, and rigid recipes should give way to guidance. AGENTS.md should explain when particular documents and tests are relevant and explicitly authorize safe routine workflows, such as running local tests with disposable data and no production access.
The guide also asks teams to put decision boundaries in writing. That means separating actions the model can take independently from those that require approval, and replacing blanket "always ask" rules with clear boundaries. For example, the model can choose how to organize a summary but should check before changing a project's scope. The definition of done should include implementing the change, running it, inspecting the result and fixing failures. For the final response, the guide recommends a short handoff covering what changed, what was checked and what still needs attention.
Features for long-running work are presented separately for the API and Codex. In the API, mid-turn steering lets developers send a correction to a working model through the Responses WebSocket API; updates are queued and do not cancel running tools or undo completed actions. Asynchronous tool calling lets the model continue unrelated work while the app runs a slower task such as tests. GPT-6.1 Sol supports multi-agent workflows in the Responses API, assigning independent work such as investigating different parts of a codebase to subagents and combining their findings into one response. Multi-agent is still in beta. In Codex, GPT-6 Astra can ask the user clarifying questions while it works, and users can redirect an active task by feeding it new information.
Computer use is open to all three models. According to the guide, GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna can interact directly with websites and desktop apps, including applications without an API. That means a single flow can cover investigating a bug, fixing the code and opening the product in a browser to confirm the fix works. OpenAI says to use an API or connected tool when it can do the job directly, and to use computer use only when the model needs to read a screen, click buttons or fill in a form. For developers building it into their own apps, the guide points to Playwright for browsers and PyAutoGUI for desktop apps.
The guide also includes four customer cases of GPT-6 Astra in production. Legal AI company Harvey combines court information, case law, firm documents and a lawyer's preferences to tailor its drafts. "We can give more context to the model and produce better and better structured outputs," said cofounder Gabe Pereyra. Cognition uses GPT-6 Astra inside Devin to test software and return evidence; in an iPhone-game example, Devin produced a simulator recording along with a report separating checks that passed from areas left untested. Hex turns questions about sales-channel performance into written findings with geographic breakdowns and interactive dashboards, and asks the model to examine whether the numbers make sense. Invideo said its success rate on color-grading tasks rose roughly threefold, and that a few editors created about 50 effects in one day.
METAL has previously reported on OpenAI's launch of GPT-6.1 Sol, and this guide is an official document in which OpenAI itself lays out the uses of the three models three days after Sol's release. The original guide METAL reviewed carries model card images along with accounts from a developer who ran a long refactor to asynchronous workers with GPT-6 Astra and another who fixed iPad compatibility using Ultrafast and live steering.
Seen from the perspective of AI engineers and technical project managers, the guide clearly shifts the center of gravity. Raising performance is no longer about writing longer prompts but about deciding what to delegate, where to stop and what counts as done. Cost management has also moved from picking a single model to a design problem of combining cache structure, reasoning effort and speed options, and the unit of measure OpenAI puts forward is not the per-token rate but the cost per successful task. As models take on work lasting days, the instruction documents that define boundaries and completion criteria are becoming as central a deliverable for development teams as the code itself.





Comments