AI news and explainers at 7 AM weekdays, plus a Sunday weekly at 8Get it in your inbox

METAL LAB

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

arXiv:2608.205742026-08-24

A benchmark that grades AI cooking decisions with a frozen scoring table instead of a human or AI judge

FlavourBench asks 27 frontier language models to pick 3 of 8 ingredients across 534 identical tasks, and a culinary system called Epicure pre-scores all 56 possible combinations before any model answers, so grading is a simple table lookup. Grok 4.6 scored highest at 65.1, but only 101 of 351 pairwise model comparisons were statistically resolved after correction for multiple testing. The release publishes every prompt, score map, raw response, and an offline verifier so the entire leaderboard can be reproduced.

METAL LAB explanatory visual

How FlavourBench evaluates and scores models

Evidence statusMeasured results reported

  1. Task generationEight-ingredient tasks are created and Epicure pre-scores all 56 possible three-ingredient picks on a frozen 0-100 scale before any model sees them
  2. Run 27 modelsAll models answer the identical 534 tasks; only tasks every model completes enter the common core, giving 14,418 model-task cells
  3. Table-lookup scoringA model's chosen three-ingredient combination is scored by looking it up in the frozen score map, with no judge model or human rater involved
  4. Statistical testing50,000 bootstrap resamples build confidence intervals and 100,000 sign-flip resamples with Holm correction test all 351 pairwise model comparisons
  5. Open verificationPrompts, score maps, raw responses, and exact routes are published with an offline verifier that reconstructs every reported result
An explanatory diagram made by METAL LAB, not a figure supplied by the paper's authors.

What they did

  1. Instead of relying on human raters or another AI model as judge, FlavourBench uses Epicure, a culinary system, to pre-score all 56 possible three-ingredient combinations for each eight-ingredient task before any model responds, so grading is a direct table lookup.
  2. It covers three task families -- ingredient substitution, pairing, and constrained composition -- totaling 534 tasks, and runs 27 frontier endpoints including GPT-5.6 variants, Claude, Gemini, Grok, Llama, Qwen, and DeepSeek on the identical set.
  3. Only tasks completed by every one of the 27 models are counted in the leaderboard (89 tasks per panel and family, 14,418 model-task cells total), removing the possibility that some models look better simply by answering easier tasks.
  4. Score confidence intervals come from 50,000 bootstrap resamples and pairwise model comparisons come from 100,000 sign-flip resamples, with Holm correction applied to control for false positives across many comparisons.
  5. Two independently compiled task panels were used to check whether the ranking replicates on a disjoint set of tasks.
Figure 2: FlavourBench Score with simultaneous 95% max-t intervals. The letter-free group numbers are inferential tiers: a point rank is shown for navigation, while models inside an unresolved group should not be read as a statistically established ordering.
Figure 2: FlavourBench Score with simultaneous 95% max-t intervals. The letter-free group numbers are inferential tiers: a point rank is shown for navigation, while models inside an unresolved group should not be read as a statistically established ordering.
Table 1: The three complete-core task families share one response and scoring interface but probe different culinary decisions. Each task contains eight candidates and 56 frozen scores.
FamilyDecisionEpicure utilityHard constraintTasks
Substitutionchoose a three-item replacement portfolio for an anchor ingredient0.8 anchor similarity +0.2 portfolio coherencenone178
Pairingchoose three ingredients to accompany an anchor0.65 anchor affinity +0.35 portfolio coherencenone178
Constraintschoose a feasible portfolio under diet and processing limits0.7 anchor similarity +0.3 coherencediet and maximum NOVA level178
Figure 3: Panel-1 versus panel-2 FlavourBench Score for every endpoint. The diagonal marks exact agreement. Each point aggregates 89 tasks per family; the joint analysis uses 534 unique anchor clusters.
Figure 3: Panel-1 versus panel-2 FlavourBench Score for every endpoint. The diagonal marks exact agreement. Each point aggregates 89 tasks per family; the joint analysis uses 534 unique anchor clusters.
Table 2: Frozen task-map diagnostics for 178 scheduled tasks per family, computed before model execution. Chance is the uniformly random portfolio mean; top gap is the median best–second-best margin.
FamilyExact chanceTop gapDistinct scores
Substitution45.25.656
Pairing45.05.656
Constraints3.437.34
Figure 4: All paired model contrasts after Holm correction. Teal means the row model scores higher; red means lower; grey is unresolved.
Figure 4: All paired model contrasts after Holm correction. Teal means the row model scores higher; red means lower; grey is unresolved.
Table 3: The complete automated leaderboard. Every score and every pairwise comparison uses the identical 534-task core. Rank intervals come from the anchor-cluster bootstrap; groups summarize the Holm-controlled pairwise graph.
RankModelFB Scoresimultaneous 95% CIrank 95% CIgroup
1Grok 4.665.1[61.0, 69.2][1, 5]1
2Gemini 3.1 Pro65.0[60.8, 69.1][1, 6]1
3GPT-5.6 Sol Pro64.2[60.1, 68.4][1, 8]1
4Muse Spark 1.263.8[59.6, 67.9][1, 10]1
5GPT-5.6 Terra Pro63.7[59.5, 67.8][1, 11]1
6Claude Fable 563.4[59.2, 67.5][2, 13]1
7GPT-5.6 Luna Pro62.6[58.5, 66.8][4, 14]1
8Claude Opus 562.5[58.4, 66.6][3, 15]1
9Qwen3.8 A95B62.1[57.9, 66.2][4, 17]1
10Kimi K362.1[57.9, 66.2][5, 16]1
11Gemini 3.6 Flash62.0[57.7, 66.2][4, 17]1
12DeepSeek V4 Pro 081362.0[57.8, 66.1][5, 17]1
13Qwen 3.8 Max61.5[57.4, 65.6][7, 18]1
14Tencent HY 361.5[57.3, 65.7][6, 19]1
15MiniMax M360.9[56.8, 65.1][8, 20]1
16GLM 5.360.6[56.4, 64.8][8, 20]1
17Muse Glimmer 30B59.9[55.9, 63.9][12, 21]2
18Seed 2.1 Turbo59.7[55.4, 64.0][12, 22]2
19Inkling59.6[55.4, 63.8][12, 22]2
20Claude Sonnet 559.5[55.4, 63.7][13, 22]2
21GLM 5.258.5[54.2, 62.7][16, 23]2
22Nemotron 3.5 Lightning57.4[53.2, 61.6][19, 25]2
23Command A56.7[52.5, 60.9][20, 25]2
24DeepSeek V4 Flash55.4[51.1, 59.7][22, 26]2
25Mistral Large 355.4[51.4, 59.5][22, 26]2
26Llama 4 Maverick53.7[49.6, 57.7][24, 26]3
27Command R+47.9[43.7, 52.0][27, 27]3
Figure 5: Mean score by family on the common core. Each family contributes 178 tasks per model. Values share the 0–100 scale but retain family-specific decision semantics.
Figure 5: Mean score by family on the common core. Each family contributes 178 tasks per model. Values share the 0–100 scale but retain family-specific decision semantics.
Table 4: Three common-core response examples. Parentheses give the fixed task score. The complete prompts and unabridged responses are released in machine-readable form.
FamilyPromptHigher-ranked selectionLower-ranked selection
SubstitutionSelect three alternatives to ‘boursin cheese‘ that best preserve its dairy role, regional context, and mutual portfolio coherence.caciocavallo, fromage blanc, grana padano (100)fromage blanc, quail egg, goose egg (15)
PairingSelect the three-ingredient bundle with the strongest learned pairing to ‘sweetbread‘ and the best internal coherence.parsnip, cipollini onion, porcini mushroom (92)parsnip, pickled onion, peppadew pepper (0)
ConstraintsFor a dish centered on ‘potato‘, select three vegetarian complements with NOVA processing level at most 2; among valid portfolios, maximize learned pairing and coherence.parsley, black pepper, thyme (100)cheese, bay leaf, parsley (0)
Figure 6: Point rank and 95% bootstrap rank interval for every endpoint. Narrow intervals imply a stable location; overlapping intervals mark unresolved local order.
Figure 6: Point rank and 95% bootstrap rank interval for every endpoint. Narrow intervals imply a stable location; overlapping intervals mark unresolved local order.

Findings

  • Across 27 models evaluated on the same 534 tasks, Grok 4.6 had the highest point estimate at 65.1, with a simultaneous 95% confidence interval of 61.0 to 69.2.
  • The two independently compiled task panels correlated at r=0.89 in model-level score and rank correlation rho=0.80.
  • Of 351 pairwise model comparisons, 101 remained statistically significant after Holm correction.
  • Only tasks completed by all models entered the leaderboard, yielding exactly 534 valid responses per model and 14,418 model-task cells total.

Where it can be used

  • The design offers a template for building reproducible, judge-free leaderboards for evaluating open-ended language model decisions.
  • The published prompts, score maps, and offline verifier let other researchers reproduce results or evaluate additional models under the identical scoring criteria.
  • The dense 56-way reward maps per task could be used to design follow-up experiments in supervised learning, preference learning, or reinforcement learning.

Limits and open work

  • Scores measure agreement with one specific, versioned release of Epicure, not universal human culinary taste.
  • The tasks cover constrained ingredient selection only, not full recipe generation, actual cooking execution, safety advice, or long-horizon kitchen planning.
  • Requiring every model to complete a task restricts the evaluated set to a common core; the excluded tasks are only released as a separate supplemental track.
  • Results are tied to specific model routes and collection dates, so they do not automatically apply to later model versions or different access paths.
  • The paper does not test whether training a model to optimize Epicure's reward actually transfers to real human cooking outcomes.

Why it matters

Judging model answers with another AI or a small human panel risks entangling the judge with the systems being tested, and this frozen-scoring-table approach avoids that by removing any judge model from the loop. It also reports which model comparisons are statistically confirmed versus unresolved rather than presenting a single ranked list, which matters for anyone tempted to read a leaderboard position as a settled fact.

Terms in this paper

  • Epicure · A culinary system representing 1,790 ingredients in a 300-dimensional space that computes substitution, pairing, and dietary feasibility, used here as the scoring reference
  • anchor-cluster bootstrap · A resampling method that groups tasks sharing the same ingredient together when repeatedly resampling to compute confidence intervals
  • Holm correction · A statistical adjustment that tightens significance thresholds when many comparisons are run at once, to avoid false positives
  • simultaneous 95% confidence interval · An interval guaranteed to contain the true value with 95% probability even when viewed across all 27 model scores at once
  • constrained composition · A task type where a dietary or category rule must be satisfied; infeasible ingredient combinations score zero

Original abstract (English)

Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two independently compiled panels correlate at r = 0.89 (rank rho = 0.80). Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% CI 61.0-69.2); 101 of 351 model pairs are resolved. The release includes the prompts, all portfolio score maps, raw responses, exact routes, content hashes, and an offline verifier that reconstructs every result.

Authors · Josef Chen, Erim Hayretci

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB

Figures: Josef Chen et al., arXiv:2608.20574, CC BY 4.0