
Summary
- Databricks unveiled a new benchmark, OfficeQA Pro V2, in an August 6, 2026 blog post.
- It is designed to evaluate "grounded reasoning" — the ability to produce answers based on enterprise materials — and the post includes agent performance results.
- Questions were first created by humans, then scaled up using synthetic data.
The questions employees typically ask AI inside a company tend to look like this: "Based on these reports, summarize why the Q3 contract renewal rate dropped." The answer doesn't live inside the model's head — it's buried in a pile of internal documents. The new benchmark Databricks released on August 6 measures exactly this capability.
What was released
In a company blog post, Databricks unveiled OfficeQA Pro V2, a benchmark designed to evaluate "grounded reasoning" in enterprise settings — the ability to reason by anchoring answers to given source material. The post covers, in order: performance results from running several agents on the benchmark, the process of building the new benchmark, the method used to scale it with synthetic data, example questions, and the benchmark's detailed specifications. The "V2" in the name signals that it succeeds an earlier version.
Why "grounded reasoning" is a hard problem
A model doing well on exam-style questions is a different thing from reading and answering based on company documents. The latter fails in messier ways. A model might plausibly fabricate a number that isn't in the source material (hallucination), or draw a conclusion from only the first of three documents when the answer is actually scattered across all three, or correctly locate the relevant passage but botch the calculation built on it. When the input mixes formats — PDF contracts, quarterly earnings tables, presentation slides — the difficulty jumps another level.
The problem is that existing benchmarks don't catch these failures well. Knowledge-test-style evaluations like MMLU ask what a model already knows. What enterprises actually want to know is whether a model performs well when handed their own documents. As a result, it has become common for companies to build their own private evaluation sets from internal materials — but since these aren't public, there's no way to compare across them. That's why public benchmarks keep emerging.
There's another wall here: these kinds of questions need to be hand-crafted by humans to be high quality, which caps the dataset at a few hundred items. The fact that Databricks devoted a separate section to "scaling with synthetic data" reads as its answer to how it got past this bottleneck. Generating questions with a model solves the scale problem, but raises new challenges — keeping difficulty consistent, and verifying that answers are actually grounded in the source material.
Why Databricks is releasing a benchmark
Databricks is a data platform company founded by the researchers who created Apache Spark. Its core business is layering analytics and AI on top of enterprise data, and in recent years it has shifted its focus toward agents running on top of that data. On August 10, it relaunched its 20-city "Data + AI World Tour," and has also been rolling out case studies of agent orchestration with accounting and advisory firm CLA. Agents that work grounded in internal company data are a core use case for the company's products. For a company like this to build its own evaluation standard is also, in effect, a statement about how it believes the problem should be defined.
What changes as a result
For organizations evaluating internal AI adoption, there's now one more reference point. It means the question "can this model work with our documents?" can be tested against a public set of questions rather than gut feeling. Viewed as a trend, this aligns with a shift in model performance competition — from "what does it know" to "what can it do with the materials it's given." Attempts to turn benchmarking itself into a product are also increasing — Artificial Analysis opened early access sign-ups for a developer-focused benchmarking product suite on August 10. That said, for specific scores on which agents performed how well in this release, readers will need to check the results table in the original post directly.





Comments