工作日早上 7 点读 AI,周日早上 8 点读周报订阅邮件

METAL LAB

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

arXiv:2608.205742026-08-24

用一套可执行的美食评分系统当裁判,而不是靠人或AI来打分的语言模型评测

FlavourBench让27个前沿语言模型在534个相同任务中,从8种食材里选出3种搭配,而一个叫Epicure的美食系统会提前给所有56种可能的组合打好分,评分只需查表即可完成。结果显示Grok 4.6得分最高,为65.1分,但在351组两两模型比较中,经过多重检验校正后只有101组结果具有统计显著性。研究团队公开了全部提示词、评分表、原始回答和离线验证工具,使整个排行榜可以被完整复现。

METAL LAB 解读图

FlavourBench的评测与打分流程

证据状态已报告实测结果

  1. 任务生成为每个八食材任务生成全部56种三食材组合,由Epicure提前打出0到100分的固定分数表
  2. 27个模型作答所有模型在相同的534个任务上作答,只有全部模型都完成的任务进入共同核心集,共14418个模型-任务格
  3. 查表打分模型选出的三食材组合直接在冻结好的分数表中查找得分,过程中没有裁判模型或人工评审参与
  4. 统计检验5万次自助重采样构建置信区间,10万次符号翻转重采样配合Holm校正检验351组两两模型比较
  5. 公开验证提示词、评分表、原始回答和调用路径全部公开,并提供离线验证工具重建所有已报告结果
这是 METAL LAB 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 不再依赖人类评审团或另一个AI模型来打分,而是用美食系统Epicure提前为每个任务的全部56种三食材组合打好0到100分的固定分数表,评分时直接查表即可,无需裁判模型参与。
  2. 任务分为食材替换、搭配、约束条件组合三类,共534个任务,27个前沿模型(包括GPT-5.6系列、Claude、Gemini、Grok、Llama、Qwen、DeepSeek等)在完全相同的任务集上接受测试。
  3. 只有全部27个模型都完成的任务才计入最终排行榜(每个模型每个面板每个家族恰好89个任务,共14,418个模型-任务格),避免某些模型因为只答了更简单的任务而显得成绩更好。
  4. 分数的置信区间通过5万次自助重采样计算,351组模型两两比较通过10万次符号翻转重采样检验,并采用Holm校正来控制多重比较带来的误报。
  5. 使用两套独立编制的任务面板来检验排名结果能否在完全不同的任务集上复现。
Figure 2: FlavourBench Score with simultaneous 95% max-t intervals. The letter-free group numbers are inferential tiers: a point rank is shown for navigation, while models inside an unresolved group should not be read as a statistically established ordering.
Figure 2: FlavourBench Score with simultaneous 95% max-t intervals. The letter-free group numbers are inferential tiers: a point rank is shown for navigation, while models inside an unresolved group should not be read as a statistically established ordering.
Table 1: The three complete-core task families share one response and scoring interface but probe different culinary decisions. Each task contains eight candidates and 56 frozen scores.
FamilyDecisionEpicure utilityHard constraintTasks
Substitutionchoose a three-item replacement portfolio for an anchor ingredient0.8 anchor similarity +0.2 portfolio coherencenone178
Pairingchoose three ingredients to accompany an anchor0.65 anchor affinity +0.35 portfolio coherencenone178
Constraintschoose a feasible portfolio under diet and processing limits0.7 anchor similarity +0.3 coherencediet and maximum NOVA level178
Figure 3: Panel-1 versus panel-2 FlavourBench Score for every endpoint. The diagonal marks exact agreement. Each point aggregates 89 tasks per family; the joint analysis uses 534 unique anchor clusters.
Figure 3: Panel-1 versus panel-2 FlavourBench Score for every endpoint. The diagonal marks exact agreement. Each point aggregates 89 tasks per family; the joint analysis uses 534 unique anchor clusters.
Table 2: Frozen task-map diagnostics for 178 scheduled tasks per family, computed before model execution. Chance is the uniformly random portfolio mean; top gap is the median best–second-best margin.
FamilyExact chanceTop gapDistinct scores
Substitution45.25.656
Pairing45.05.656
Constraints3.437.34
Figure 4: All paired model contrasts after Holm correction. Teal means the row model scores higher; red means lower; grey is unresolved.
Figure 4: All paired model contrasts after Holm correction. Teal means the row model scores higher; red means lower; grey is unresolved.
Table 3: The complete automated leaderboard. Every score and every pairwise comparison uses the identical 534-task core. Rank intervals come from the anchor-cluster bootstrap; groups summarize the Holm-controlled pairwise graph.
RankModelFB Scoresimultaneous 95% CIrank 95% CIgroup
1Grok 4.665.1[61.0, 69.2][1, 5]1
2Gemini 3.1 Pro65.0[60.8, 69.1][1, 6]1
3GPT-5.6 Sol Pro64.2[60.1, 68.4][1, 8]1
4Muse Spark 1.263.8[59.6, 67.9][1, 10]1
5GPT-5.6 Terra Pro63.7[59.5, 67.8][1, 11]1
6Claude Fable 563.4[59.2, 67.5][2, 13]1
7GPT-5.6 Luna Pro62.6[58.5, 66.8][4, 14]1
8Claude Opus 562.5[58.4, 66.6][3, 15]1
9Qwen3.8 A95B62.1[57.9, 66.2][4, 17]1
10Kimi K362.1[57.9, 66.2][5, 16]1
11Gemini 3.6 Flash62.0[57.7, 66.2][4, 17]1
12DeepSeek V4 Pro 081362.0[57.8, 66.1][5, 17]1
13Qwen 3.8 Max61.5[57.4, 65.6][7, 18]1
14Tencent HY 361.5[57.3, 65.7][6, 19]1
15MiniMax M360.9[56.8, 65.1][8, 20]1
16GLM 5.360.6[56.4, 64.8][8, 20]1
17Muse Glimmer 30B59.9[55.9, 63.9][12, 21]2
18Seed 2.1 Turbo59.7[55.4, 64.0][12, 22]2
19Inkling59.6[55.4, 63.8][12, 22]2
20Claude Sonnet 559.5[55.4, 63.7][13, 22]2
21GLM 5.258.5[54.2, 62.7][16, 23]2
22Nemotron 3.5 Lightning57.4[53.2, 61.6][19, 25]2
23Command A56.7[52.5, 60.9][20, 25]2
24DeepSeek V4 Flash55.4[51.1, 59.7][22, 26]2
25Mistral Large 355.4[51.4, 59.5][22, 26]2
26Llama 4 Maverick53.7[49.6, 57.7][24, 26]3
27Command R+47.9[43.7, 52.0][27, 27]3
Figure 5: Mean score by family on the common core. Each family contributes 178 tasks per model. Values share the 0–100 scale but retain family-specific decision semantics.
Figure 5: Mean score by family on the common core. Each family contributes 178 tasks per model. Values share the 0–100 scale but retain family-specific decision semantics.
Table 4: Three common-core response examples. Parentheses give the fixed task score. The complete prompts and unabridged responses are released in machine-readable form.
FamilyPromptHigher-ranked selectionLower-ranked selection
SubstitutionSelect three alternatives to ‘boursin cheese‘ that best preserve its dairy role, regional context, and mutual portfolio coherence.caciocavallo, fromage blanc, grana padano (100)fromage blanc, quail egg, goose egg (15)
PairingSelect the three-ingredient bundle with the strongest learned pairing to ‘sweetbread‘ and the best internal coherence.parsnip, cipollini onion, porcini mushroom (92)parsnip, pickled onion, peppadew pepper (0)
ConstraintsFor a dish centered on ‘potato‘, select three vegetarian complements with NOVA processing level at most 2; among valid portfolios, maximize learned pairing and coherence.parsley, black pepper, thyme (100)cheese, bay leaf, parsley (0)
Figure 6: Point rank and 95% bootstrap rank interval for every endpoint. Narrow intervals imply a stable location; overlapping intervals mark unresolved local order.
Figure 6: Point rank and 95% bootstrap rank interval for every endpoint. Narrow intervals imply a stable location; overlapping intervals mark unresolved local order.

研究结果

  • 在相同的534个任务上评测27个模型后,Grok 4.6的估计分数最高,为65.1,同时95%置信区间为61.0至69.2。
  • 两套独立编制的任务面板在模型层面得分的相关系数为0.89,排名相关系数为0.80。
  • 在351组两两模型比较中,经Holm校正后仍具统计显著性的有101组。
  • 排行榜只纳入全部模型都完成的任务,每个模型恰好获得534个有效回答,总计14,418个模型-任务格。

可应用场景

  • 该设计为构建无需裁判模型、可复现的开放式语言模型评测排行榜提供了一种参考思路。
  • 公开的提示词、评分表和离线验证工具可供其他研究者复现结果,或在相同标准下评测新增模型。
  • 每个任务56种组合的密集奖励分数表,可用于设计后续的监督学习、偏好学习或强化学习实验。

局限与待验证事项

  • 分数衡量的是与某一特定版本的Epicure系统的一致程度,并不代表人类普遍的美食品味。
  • 评测任务局限于有限的食材选择决策,不涉及完整菜谱生成、实际烹饪操作、安全建议或长期厨房规划。
  • 要求所有模型都完成才能计入排行榜的做法把评测范围限定在这个共同任务集合内,未纳入的任务仅作为补充数据集单独发布。
  • 结果与特定时间点、特定调用路径下的模型版本绑定,不能直接推及后续模型更新或其他调用方式。
  • 论文并未检验针对Epicure奖励进行优化训练是否真的能迁移到人类实际烹饪效果上。

为什么重要

用另一个AI模型或少量人工评审来打分,容易让裁判本身与被测系统产生纠缠或偏见,而这种预先冻结的评分表方法从流程中彻底去掉了裁判模型。它还明确区分了哪些模型间差异经统计检验确认、哪些尚未确认,提醒读者不要把排行榜的名次简单当作已经证实的定论。

本文术语

  • Epicure · 一个将1790种食材表示为300维空间的美食系统,可计算食材替换、搭配和饮食可行性,在本研究中充当评分的参照标准
  • 锚点聚类自助法 · 一种重采样统计方法,把共享同一食材的任务捆绑在一起反复重采样,用来计算置信区间
  • Holm校正 · 在同时进行多次统计比较时收紧显著性标准的一种方法,用来避免偶然出现的假阳性结果
  • 同时95%置信区间 · 即使同时查看全部27个模型的分数,该区间也能以95%的概率包含真实值
  • 约束条件组合任务 · 一种要求满足饮食限制等条件的任务类型,不满足条件的食材组合直接记为零分

论文原文摘要(英文)

Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two independently compiled panels correlate at r = 0.89 (rank rho = 0.80). Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% CI 61.0-69.2); 101 of 351 model pairs are resolved. The release includes the prompts, all portfolio score maps, raw responses, exact routes, content hashes, and an offline verifier that reconstructs every result.

作者 · Josef Chen, Erim Hayretci

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道

图片来源: Josef Chen et al., arXiv:2608.20574, CC BY 4.0