FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
用一套可执行的美食评分系统当裁判,而不是靠人或AI来打分的语言模型评测
FlavourBench让27个前沿语言模型在534个相同任务中,从8种食材里选出3种搭配,而一个叫Epicure的美食系统会提前给所有56种可能的组合打好分,评分只需查表即可完成。结果显示Grok 4.6得分最高,为65.1分,但在351组两两模型比较中,经过多重检验校正后只有101组结果具有统计显著性。研究团队公开了全部提示词、评分表、原始回答和离线验证工具,使整个排行榜可以被完整复现。
METAL LAB 解读图
FlavourBench的评测与打分流程
证据状态已报告实测结果
- 任务生成为每个八食材任务生成全部56种三食材组合,由Epicure提前打出0到100分的固定分数表
- 27个模型作答所有模型在相同的534个任务上作答,只有全部模型都完成的任务进入共同核心集,共14418个模型-任务格
- 查表打分模型选出的三食材组合直接在冻结好的分数表中查找得分,过程中没有裁判模型或人工评审参与
- 统计检验5万次自助重采样构建置信区间,10万次符号翻转重采样配合Holm校正检验351组两两模型比较
- 公开验证提示词、评分表、原始回答和调用路径全部公开,并提供离线验证工具重建所有已报告结果
他们做了什么
- 不再依赖人类评审团或另一个AI模型来打分,而是用美食系统Epicure提前为每个任务的全部56种三食材组合打好0到100分的固定分数表,评分时直接查表即可,无需裁判模型参与。
- 任务分为食材替换、搭配、约束条件组合三类,共534个任务,27个前沿模型(包括GPT-5.6系列、Claude、Gemini、Grok、Llama、Qwen、DeepSeek等)在完全相同的任务集上接受测试。
- 只有全部27个模型都完成的任务才计入最终排行榜(每个模型每个面板每个家族恰好89个任务,共14,418个模型-任务格),避免某些模型因为只答了更简单的任务而显得成绩更好。
- 分数的置信区间通过5万次自助重采样计算,351组模型两两比较通过10万次符号翻转重采样检验,并采用Holm校正来控制多重比较带来的误报。
- 使用两套独立编制的任务面板来检验排名结果能否在完全不同的任务集上复现。
| Family | Decision | Epicure utility | Hard constraint | Tasks |
|---|---|---|---|---|
| Substitution | choose a three-item replacement portfolio for an anchor ingredient | 0.8 anchor similarity +0.2 portfolio coherence | none | 178 |
| Pairing | choose three ingredients to accompany an anchor | 0.65 anchor affinity +0.35 portfolio coherence | none | 178 |
| Constraints | choose a feasible portfolio under diet and processing limits | 0.7 anchor similarity +0.3 coherence | diet and maximum NOVA level | 178 |
| Family | Exact chance | Top gap | Distinct scores |
|---|---|---|---|
| Substitution | 45.2 | 5.6 | 56 |
| Pairing | 45.0 | 5.6 | 56 |
| Constraints | 3.4 | 37.3 | 4 |

| Rank | Model | FB Score | simultaneous 95% CI | rank 95% CI | group |
|---|---|---|---|---|---|
| 1 | Grok 4.6 | 65.1 | [61.0, 69.2] | [1, 5] | 1 |
| 2 | Gemini 3.1 Pro | 65.0 | [60.8, 69.1] | [1, 6] | 1 |
| 3 | GPT-5.6 Sol Pro | 64.2 | [60.1, 68.4] | [1, 8] | 1 |
| 4 | Muse Spark 1.2 | 63.8 | [59.6, 67.9] | [1, 10] | 1 |
| 5 | GPT-5.6 Terra Pro | 63.7 | [59.5, 67.8] | [1, 11] | 1 |
| 6 | Claude Fable 5 | 63.4 | [59.2, 67.5] | [2, 13] | 1 |
| 7 | GPT-5.6 Luna Pro | 62.6 | [58.5, 66.8] | [4, 14] | 1 |
| 8 | Claude Opus 5 | 62.5 | [58.4, 66.6] | [3, 15] | 1 |
| 9 | Qwen3.8 A95B | 62.1 | [57.9, 66.2] | [4, 17] | 1 |
| 10 | Kimi K3 | 62.1 | [57.9, 66.2] | [5, 16] | 1 |
| 11 | Gemini 3.6 Flash | 62.0 | [57.7, 66.2] | [4, 17] | 1 |
| 12 | DeepSeek V4 Pro 0813 | 62.0 | [57.8, 66.1] | [5, 17] | 1 |
| 13 | Qwen 3.8 Max | 61.5 | [57.4, 65.6] | [7, 18] | 1 |
| 14 | Tencent HY 3 | 61.5 | [57.3, 65.7] | [6, 19] | 1 |
| 15 | MiniMax M3 | 60.9 | [56.8, 65.1] | [8, 20] | 1 |
| 16 | GLM 5.3 | 60.6 | [56.4, 64.8] | [8, 20] | 1 |
| 17 | Muse Glimmer 30B | 59.9 | [55.9, 63.9] | [12, 21] | 2 |
| 18 | Seed 2.1 Turbo | 59.7 | [55.4, 64.0] | [12, 22] | 2 |
| 19 | Inkling | 59.6 | [55.4, 63.8] | [12, 22] | 2 |
| 20 | Claude Sonnet 5 | 59.5 | [55.4, 63.7] | [13, 22] | 2 |
| 21 | GLM 5.2 | 58.5 | [54.2, 62.7] | [16, 23] | 2 |
| 22 | Nemotron 3.5 Lightning | 57.4 | [53.2, 61.6] | [19, 25] | 2 |
| 23 | Command A | 56.7 | [52.5, 60.9] | [20, 25] | 2 |
| 24 | DeepSeek V4 Flash | 55.4 | [51.1, 59.7] | [22, 26] | 2 |
| 25 | Mistral Large 3 | 55.4 | [51.4, 59.5] | [22, 26] | 2 |
| 26 | Llama 4 Maverick | 53.7 | [49.6, 57.7] | [24, 26] | 3 |
| 27 | Command R+ | 47.9 | [43.7, 52.0] | [27, 27] | 3 |

| Family | Prompt | Higher-ranked selection | Lower-ranked selection |
|---|---|---|---|
| Substitution | Select three alternatives to ‘boursin cheese‘ that best preserve its dairy role, regional context, and mutual portfolio coherence. | caciocavallo, fromage blanc, grana padano (100) | fromage blanc, quail egg, goose egg (15) |
| Pairing | Select the three-ingredient bundle with the strongest learned pairing to ‘sweetbread‘ and the best internal coherence. | parsnip, cipollini onion, porcini mushroom (92) | parsnip, pickled onion, peppadew pepper (0) |
| Constraints | For a dish centered on ‘potato‘, select three vegetarian complements with NOVA processing level at most 2; among valid portfolios, maximize learned pairing and coherence. | parsley, black pepper, thyme (100) | cheese, bay leaf, parsley (0) |
研究结果
- 在相同的534个任务上评测27个模型后,Grok 4.6的估计分数最高,为65.1,同时95%置信区间为61.0至69.2。
- 两套独立编制的任务面板在模型层面得分的相关系数为0.89,排名相关系数为0.80。
- 在351组两两模型比较中,经Holm校正后仍具统计显著性的有101组。
- 排行榜只纳入全部模型都完成的任务,每个模型恰好获得534个有效回答,总计14,418个模型-任务格。
可应用场景
- 该设计为构建无需裁判模型、可复现的开放式语言模型评测排行榜提供了一种参考思路。
- 公开的提示词、评分表和离线验证工具可供其他研究者复现结果,或在相同标准下评测新增模型。
- 每个任务56种组合的密集奖励分数表,可用于设计后续的监督学习、偏好学习或强化学习实验。
局限与待验证事项
- 分数衡量的是与某一特定版本的Epicure系统的一致程度,并不代表人类普遍的美食品味。
- 评测任务局限于有限的食材选择决策,不涉及完整菜谱生成、实际烹饪操作、安全建议或长期厨房规划。
- 要求所有模型都完成才能计入排行榜的做法把评测范围限定在这个共同任务集合内,未纳入的任务仅作为补充数据集单独发布。
- 结果与特定时间点、特定调用路径下的模型版本绑定,不能直接推及后续模型更新或其他调用方式。
- 论文并未检验针对Epicure奖励进行优化训练是否真的能迁移到人类实际烹饪效果上。
为什么重要
用另一个AI模型或少量人工评审来打分,容易让裁判本身与被测系统产生纠缠或偏见,而这种预先冻结的评分表方法从流程中彻底去掉了裁判模型。它还明确区分了哪些模型间差异经统计检验确认、哪些尚未确认,提醒读者不要把排行榜的名次简单当作已经证实的定论。
本文术语
- Epicure · 一个将1790种食材表示为300维空间的美食系统,可计算食材替换、搭配和饮食可行性,在本研究中充当评分的参照标准
- 锚点聚类自助法 · 一种重采样统计方法,把共享同一食材的任务捆绑在一起反复重采样,用来计算置信区间
- Holm校正 · 在同时进行多次统计比较时收紧显著性标准的一种方法,用来避免偶然出现的假阳性结果
- 同时95%置信区间 · 即使同时查看全部27个模型的分数,该区间也能以95%的概率包含真实值
- 约束条件组合任务 · 一种要求满足饮食限制等条件的任务类型,不满足条件的食材组合直接记为零分
论文原文摘要(英文)
Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two independently compiled panels correlate at r = 0.89 (rank rho = 0.80). Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% CI 61.0-69.2); 101 of 351 model pairs are resolved. The release includes the prompts, all portfolio score maps, raw responses, exact routes, content hashes, and an offline verifier that reconstructs every result.
在 arXiv 阅读最新论文
- AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scaleAI不再为单个任务搭环境,而是直接生成整个商业世界,让训练场自己扩展规模
- When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha治疗聊天机器人能听懂青少年的俚语,却常常判断不出话里藏着的危机
- Hadith computational science in the age of large language models: a critical narrative review一篇批判性综述追问:大语言模型给圣训(hadith)计算研究带来的进步哪些是真的,哪些只是在窄基准上好看
- Lost in Translation: How Universal Ethical Values Fail to Translate Across Global Contexts对10个国家14位AI专家的访谈发现,'公平''透明'这些价值观在不同地方的理解并不相同
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
METAL LAB 最新报道
图片来源: Josef Chen et al., arXiv:2608.20574, CC BY 4.0