
Summary
- On September 28, OpenAI published a case study on how AI accounting startup Basis adopted GPT-6 Astra.
- In Basis's test, Astra completed a 50-tab tax workbook in half the time GPT-5.6 Sol took, and internal evaluation scores rose about 20%.
- Basis said it adjusts reasoning effort to each step's difficulty while keeping the cache intact, cutting the cost and response time of long-running tasks.
OpenAI published a case study on September 28 on how AI accounting startup Basis adopted GPT-6 Astra. When Basis gave a complex tax workbook made up of 50 tabs to GPT-6 Astra and to the earlier model GPT-5.6 Sol, Astra finished the job in half the time Sol needed. Basis's internal evaluation scores also rose about 20%. "GPT-6 Astra does a better job of really understanding the intent of the user and the problem," said co-founder Mitch Troyanovsky.
Basis automates much of the manual work accountants do each day with AI agents. Matthew Harpe, its CEO, and Troyanovsky founded the company in New York in 2023, and on February 24 it raised a $100 million Series B led by Accel at a valuation of $1.15 billion. GV, former Goldman Sachs CEO Lloyd Blankfein and existing investor Khosla Ventures took part, bringing total funding to $138 million. According to the company, about 30% of the top 25 U.S. accounting firms use Basis, and according to reports the share is 20% among the top 150 firms.
According to the OpenAI case study METAL reviewed, the task in the test was to complete the 50-tab tax workbook accurately and reliably from start to finish. "GPT-6 Astra is able to complete that workbook in half the time that GPT-5.6 Sol is able to," Troyanovsky said. He pointed to decisions made early in the task as the source of the speed. Because Astra makes better decisions from the start, the agents take a more direct path through the work and spend less time correcting mistakes, which also means they use fewer tokens. In the video he said that even with a complicated piece of work, deciding well up front on how to do it lets you finish in effectively half the time.
From an engineer's point of view, the most notable part of the case is reasoning-effort control. As a task progresses, Basis changes how much reasoning Astra uses depending on how hard each step is. It dials computation up for difficult steps and down for easier ones, and the cache stays intact while it makes these adjustments. If changing reasoning settings mid-task meant discarding the cache built up so far, cost and response time would balloon again, so this means Astra can be tuned without that penalty. Troyanovsky said this makes long-running tasks more economical for Basis and its customers.
The evaluation method does not look at output alone. Basis grades not only the agents' final answers but how they work, checking whether they follow templates, consult primary sources for tax questions and check their own work. The company explained that the roughly 20% gain in internal evaluation scores came from better reading of user intent, including knowing when to ask questions, flagging assumptions and following instructions. Astra inferred these expectations from a broader context even with fewer explicit instructions, reducing the need for Basis to write rules for individual situations.
The numbers in the video make the point clearer. Troyanovsky said internal evaluations can cover a thousand situations but agents in production face a hundred thousand, so Basis needs to be confident that performance generalizes from the thousand to the hundred thousand. Basis said that writing fewer situation-specific rules has given it more confidence that its agents can handle situations beyond those covered in internal tests. That is also why he defines the company's research as long-horizon agents that must be reliable at production scale.
Basis has focused on long-horizon agents that carry out complex accounting work over hours or even days. The company said it demonstrated the first AI agent to autonomously complete, end to end, a U.S. partnership tax return (Form 1065), which requires producing individualized tax documents for every partner and tracking custom profit-sharing arrangements. Investor Vinod Khosla, founder of Khosla Ventures, said, "Basis is already radically changing how work gets done at the best firms, driving 20% to 50% efficiencies across practices."
The accountant's role is changing too. In the video, Troyanovsky said Basis turns accountants from doers of the work into reviewers of the work. According to reports, Harpe said the goal is to ease the U.S. accountant shortage and automate routine tax returns and financial statements so accountants can focus on helping clients make key decisions such as tax strategy and capital allocation. At the time of the funding announcement, Troyanovsky also said that coding agents writing a lot of code does not mean hiring fewer engineers, and that the company is in fact hiring more.
OpenAI has been releasing industry-specific customer cases one after another since launching GPT-6 Astra. METAL previously reported that OpenAI published legal AI company Harvey's GPT-6 Astra case. From law to accounting, the common message is that in long tasks, early judgment and understanding of intent bring down both speed and cost. Whether that promise holds in work like accounting, where mistakes carry immediate liability, will ultimately be decided in the hundred thousand real-world situations that evaluations cannot cover.





Comments