FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
A benchmark that grades AI cooking decisions with a frozen scoring table instead of a human or AI judge
FlavourBench asks 27 frontier language models to pick 3 of 8 ingredients across 534 identical tasks, and a culinary system called Epicure pre-scores all 56 possible combinations before any model answers, so grading is a simple table lookup. Grok 4.6 scored highest at 65.1, but only 101 of 351 pairwise model comparisons were statistically resolved after correction for multiple testing. The release publishes every prompt, score map, raw response, and an offline verifier so the entire leaderboard can be reproduced.
METAL LAB explanatory visual
How FlavourBench evaluates and scores models
Evidence statusMeasured results reported
- Task generationEight-ingredient tasks are created and Epicure pre-scores all 56 possible three-ingredient picks on a frozen 0-100 scale before any model sees them
- Run 27 modelsAll models answer the identical 534 tasks; only tasks every model completes enter the common core, giving 14,418 model-task cells
- Table-lookup scoringA model's chosen three-ingredient combination is scored by looking it up in the frozen score map, with no judge model or human rater involved
- Statistical testing50,000 bootstrap resamples build confidence intervals and 100,000 sign-flip resamples with Holm correction test all 351 pairwise model comparisons
- Open verificationPrompts, score maps, raw responses, and exact routes are published with an offline verifier that reconstructs every reported result
What they did
- Instead of relying on human raters or another AI model as judge, FlavourBench uses Epicure, a culinary system, to pre-score all 56 possible three-ingredient combinations for each eight-ingredient task before any model responds, so grading is a direct table lookup.
- It covers three task families -- ingredient substitution, pairing, and constrained composition -- totaling 534 tasks, and runs 27 frontier endpoints including GPT-5.6 variants, Claude, Gemini, Grok, Llama, Qwen, and DeepSeek on the identical set.
- Only tasks completed by every one of the 27 models are counted in the leaderboard (89 tasks per panel and family, 14,418 model-task cells total), removing the possibility that some models look better simply by answering easier tasks.
- Score confidence intervals come from 50,000 bootstrap resamples and pairwise model comparisons come from 100,000 sign-flip resamples, with Holm correction applied to control for false positives across many comparisons.
- Two independently compiled task panels were used to check whether the ranking replicates on a disjoint set of tasks.
| Family | Decision | Epicure utility | Hard constraint | Tasks |
|---|---|---|---|---|
| Substitution | choose a three-item replacement portfolio for an anchor ingredient | 0.8 anchor similarity +0.2 portfolio coherence | none | 178 |
| Pairing | choose three ingredients to accompany an anchor | 0.65 anchor affinity +0.35 portfolio coherence | none | 178 |
| Constraints | choose a feasible portfolio under diet and processing limits | 0.7 anchor similarity +0.3 coherence | diet and maximum NOVA level | 178 |
| Family | Exact chance | Top gap | Distinct scores |
|---|---|---|---|
| Substitution | 45.2 | 5.6 | 56 |
| Pairing | 45.0 | 5.6 | 56 |
| Constraints | 3.4 | 37.3 | 4 |

| Rank | Model | FB Score | simultaneous 95% CI | rank 95% CI | group |
|---|---|---|---|---|---|
| 1 | Grok 4.6 | 65.1 | [61.0, 69.2] | [1, 5] | 1 |
| 2 | Gemini 3.1 Pro | 65.0 | [60.8, 69.1] | [1, 6] | 1 |
| 3 | GPT-5.6 Sol Pro | 64.2 | [60.1, 68.4] | [1, 8] | 1 |
| 4 | Muse Spark 1.2 | 63.8 | [59.6, 67.9] | [1, 10] | 1 |
| 5 | GPT-5.6 Terra Pro | 63.7 | [59.5, 67.8] | [1, 11] | 1 |
| 6 | Claude Fable 5 | 63.4 | [59.2, 67.5] | [2, 13] | 1 |
| 7 | GPT-5.6 Luna Pro | 62.6 | [58.5, 66.8] | [4, 14] | 1 |
| 8 | Claude Opus 5 | 62.5 | [58.4, 66.6] | [3, 15] | 1 |
| 9 | Qwen3.8 A95B | 62.1 | [57.9, 66.2] | [4, 17] | 1 |
| 10 | Kimi K3 | 62.1 | [57.9, 66.2] | [5, 16] | 1 |
| 11 | Gemini 3.6 Flash | 62.0 | [57.7, 66.2] | [4, 17] | 1 |
| 12 | DeepSeek V4 Pro 0813 | 62.0 | [57.8, 66.1] | [5, 17] | 1 |
| 13 | Qwen 3.8 Max | 61.5 | [57.4, 65.6] | [7, 18] | 1 |
| 14 | Tencent HY 3 | 61.5 | [57.3, 65.7] | [6, 19] | 1 |
| 15 | MiniMax M3 | 60.9 | [56.8, 65.1] | [8, 20] | 1 |
| 16 | GLM 5.3 | 60.6 | [56.4, 64.8] | [8, 20] | 1 |
| 17 | Muse Glimmer 30B | 59.9 | [55.9, 63.9] | [12, 21] | 2 |
| 18 | Seed 2.1 Turbo | 59.7 | [55.4, 64.0] | [12, 22] | 2 |
| 19 | Inkling | 59.6 | [55.4, 63.8] | [12, 22] | 2 |
| 20 | Claude Sonnet 5 | 59.5 | [55.4, 63.7] | [13, 22] | 2 |
| 21 | GLM 5.2 | 58.5 | [54.2, 62.7] | [16, 23] | 2 |
| 22 | Nemotron 3.5 Lightning | 57.4 | [53.2, 61.6] | [19, 25] | 2 |
| 23 | Command A | 56.7 | [52.5, 60.9] | [20, 25] | 2 |
| 24 | DeepSeek V4 Flash | 55.4 | [51.1, 59.7] | [22, 26] | 2 |
| 25 | Mistral Large 3 | 55.4 | [51.4, 59.5] | [22, 26] | 2 |
| 26 | Llama 4 Maverick | 53.7 | [49.6, 57.7] | [24, 26] | 3 |
| 27 | Command R+ | 47.9 | [43.7, 52.0] | [27, 27] | 3 |

| Family | Prompt | Higher-ranked selection | Lower-ranked selection |
|---|---|---|---|
| Substitution | Select three alternatives to ‘boursin cheese‘ that best preserve its dairy role, regional context, and mutual portfolio coherence. | caciocavallo, fromage blanc, grana padano (100) | fromage blanc, quail egg, goose egg (15) |
| Pairing | Select the three-ingredient bundle with the strongest learned pairing to ‘sweetbread‘ and the best internal coherence. | parsnip, cipollini onion, porcini mushroom (92) | parsnip, pickled onion, peppadew pepper (0) |
| Constraints | For a dish centered on ‘potato‘, select three vegetarian complements with NOVA processing level at most 2; among valid portfolios, maximize learned pairing and coherence. | parsley, black pepper, thyme (100) | cheese, bay leaf, parsley (0) |
Findings
- Across 27 models evaluated on the same 534 tasks, Grok 4.6 had the highest point estimate at 65.1, with a simultaneous 95% confidence interval of 61.0 to 69.2.
- The two independently compiled task panels correlated at r=0.89 in model-level score and rank correlation rho=0.80.
- Of 351 pairwise model comparisons, 101 remained statistically significant after Holm correction.
- Only tasks completed by all models entered the leaderboard, yielding exactly 534 valid responses per model and 14,418 model-task cells total.
Where it can be used
- The design offers a template for building reproducible, judge-free leaderboards for evaluating open-ended language model decisions.
- The published prompts, score maps, and offline verifier let other researchers reproduce results or evaluate additional models under the identical scoring criteria.
- The dense 56-way reward maps per task could be used to design follow-up experiments in supervised learning, preference learning, or reinforcement learning.
Limits and open work
- Scores measure agreement with one specific, versioned release of Epicure, not universal human culinary taste.
- The tasks cover constrained ingredient selection only, not full recipe generation, actual cooking execution, safety advice, or long-horizon kitchen planning.
- Requiring every model to complete a task restricts the evaluated set to a common core; the excluded tasks are only released as a separate supplemental track.
- Results are tied to specific model routes and collection dates, so they do not automatically apply to later model versions or different access paths.
- The paper does not test whether training a model to optimize Epicure's reward actually transfers to real human cooking outcomes.
Why it matters
Judging model answers with another AI or a small human panel risks entangling the judge with the systems being tested, and this frozen-scoring-table approach avoids that by removing any judge model from the loop. It also reports which model comparisons are statistically confirmed versus unresolved rather than presenting a single ranked list, which matters for anyone tempted to read a leaderboard position as a settled fact.
Terms in this paper
- Epicure · A culinary system representing 1,790 ingredients in a 300-dimensional space that computes substitution, pairing, and dietary feasibility, used here as the scoring reference
- anchor-cluster bootstrap · A resampling method that groups tasks sharing the same ingredient together when repeatedly resampling to compute confidence intervals
- Holm correction · A statistical adjustment that tightens significance thresholds when many comparisons are run at once, to avoid false positives
- simultaneous 95% confidence interval · An interval guaranteed to contain the true value with 95% probability even when viewed across all 27 model scores at once
- constrained composition · A task type where a dietary or category rule must be satisfied; infeasible ingredient combinations score zero
Original abstract (English)
Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two independently compiled panels correlate at r = 0.89 (rank rho = 0.80). Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% CI 61.0-69.2); 101 of 351 model pairs are resolved. The release includes the prompts, all portfolio score maps, raw responses, exact routes, content hashes, and an offline verifier that reconstructs every result.
Read on arXivLatest papers
- AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scaleAn AI system builds whole business worlds instead of single tasks, so training grounds can scale on their own
- When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation AlphaTherapy chatbots understand teen slang but still miss the crisis hidden inside it
- Hadith computational science in the age of large language models: a critical narrative reviewA critical review asks which AI advances in hadith studies are real progress and which just look good on narrow benchmarks
- Lost in Translation: How Universal Ethical Values Fail to Translate Across Global ContextsInterviews with 14 AI experts across 10 countries show that 'fairness' and 'transparency' mean different things in different places
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
Latest from METAL LAB
- Claude web and desktop now stream long answers 4x smoother
- Hugging Face reportedly courted for $13B+ buyout, but founder sounds unconvinced
- Mistral and Saudi Arabia's HUMAIN to Co-Develop Arabic AI Models
- Liquid AI open-sources on-device benchmark Pipette
- GPT-5.6 integration cuts per-task cost 82% in coding tool Kiro
Figures: Josef Chen et al., arXiv:2608.20574, CC BY 4.0