
Image: METAL
Summary
- According to a case study OpenAI published on October 8, LegalOn Technologies cut estimated daily Codex costs by about 65% compared with using GPT-5.5 alone, while maintaining development speed.
- LegalOn starts with the lightweight GPT-6 Luna, moves standard design work to GPT-6.1 Sol and escalates advanced judgment and agent orchestration to GPT-6 Astra, and it turned off Fast mode by default.
- Mature businesses were asked for up to about 20% greater efficiency while new businesses received generous budgets, and the company is building a metric that measures AI return on investment per feature release.
OpenAI on October 8 published a case study on how LegalOn Technologies cut its Codex costs. LegalOn chose among GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna according to task difficulty and allocated budgets by business stage, reducing estimated daily costs by about 65%. According to OpenAI, development speed was maintained, and mature business areas were set a goal of improving cost efficiency by up to about 20%.
LegalOn provides "Professional AI" globally, using AI to support legal work and other core business functions. OpenAI classifies the company as a startup in the Asia-Pacific region. Aiming to make the organization itself AI-native, not just its products, LegalOn brought Codex into its development process, and adoption spread across the company as more people used it day to day. That created a new problem. Unlimited use of high-performance models would drive up spending, while blanket restrictions risked undermining the productivity gained from AI.
At first LegalOn gave developers unlimited access to its main model at the time, GPT-5.5, in Fast mode. As usage expanded to design, implementation and everyday work, teams learned through experimentation how to divide work between people and AI. But staying on that path would clearly exceed the annual budget. AID CoE, the company's AI-powered development center of excellence, began drafting model-selection guidelines. AID CoE tested models and monitored usage, managers passed its findings to their teams, and engineers came to choose the right model for each task on their own.
The principle of the new guidelines is to start with a lightweight model and move up only as work gets more complex. Internal testing made the criteria more specific. The lightest model, GPT-6 Luna, handles code implementation with clear requirements and everyday automation, mainly as a subagent. GPT-6.1 Sol handles standard design, data analysis and document preparation, along with tasks that need shorter completion times than Luna. The most capable, GPT-6 Astra, handles advanced judgment such as architecture design and acts as the orchestrator directing multiple agents.
There are two layers of control. Administrator settings put monthly usage limits on departments and individuals, and AID CoE adjusts those limits as it watches usage. Fast mode was also blocked by default and opened only on individual request. Concerns arose that development would slow, but OpenAI said teams preserved performance by running tasks in parallel. Budget caps were set at three levels: department, group and individual.

Budgets were deliberately tilted by business stage. The established LegalOn business was asked to improve cost efficiency by up to about 20%, while new businesses in their launch phase received generous budgets to encourage active use of AI. Yuta Tokitake, Senior Engineering Manager at LegalOn Technologies, explained that new businesses put business speed ahead of cost efficiency. The idea is to use AI heavily to raise output and ultimately drive business growth. Combining model selection, feature restrictions and business-specific budgets cut estimated daily costs by about 65% compared with the period when only GPT-5.5 was used.
LegalOn's next question is about returns, not costs. "It is almost a given that AI can accelerate system development and updates," Tokitake said. "What we really want to understand is whether that faster development actually translates into value for customers." The company is therefore building its own metric that evaluates AI investment by customer value rather than development speed or usage volume. "When using AI, we can track the cost of individual tasks, but the total cost of a complete piece of work—a feature release—often remains a black box," he said, adding that the company designed the measurement around each feature release so it can link the customer value a release delivers to the AI costs actually invested in it. The pipeline to calculate the metric is now being built.
The OpenAI case study page that METAL reviewed includes a table dividing the roles of the three models, along with a 7-minute-12-second audio narration. METAL has previously reported on OpenAI unveiling GPT-6 Sol and Luna and the launch of GPT-6.1 Sol, as well as the Harvey GPT-6 Astra case study from the same legal field. As a next step, LegalOn plans to build an internal knowledge base that collects hands-on experience such as the best model combinations for design, implementation and review. It has already begun emphasizing AI skills in hiring, and AID CoE and the security team are jointly balancing performance, cost and risk.
From an engineering-management perspective, what this case shows is that model routing is now a budget question for development organizations. Simply switching the default from the most expensive model to the cheapest one, with escalation upward, changed the cost structure. "Excessive restrictions through rules and budgets can undermine an organization's momentum," Tokitake said. "What we need is a flexible operating model that effectively balances risk control with the speed teams need."





Comments