월~금 오전 7시, 일요일 오전 8시 — AI 뉴스와 용어를 보내드립니다메일로 받아보기

METAL LAB

AI가 '요리 결정'을 얼마나 잘하는지, 다른 AI가 아니라 실행 가능한 요리 시스템으로 채점한 벤치마크

arXiv:2608.205742026-08-24

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

AI가 '요리 결정'을 얼마나 잘하는지, 다른 AI가 아니라 실행 가능한 요리 시스템으로 채점한 벤치마크

FlavourBench는 8개 재료 중 3개를 고르는 요리 문제를 만들고, Epicure라는 요리 시스템이 가능한 56가지 조합 전부에 미리 점수를 매겨 채점 기준으로 삼는다. 27개 최신 언어모델을 동일한 534개 문제로 평가했고, Grok 4.6이 65.1점으로 가장 높았지만 351개 모델 짝 비교 중 101개만 통계적으로 순위가 확정됐다. 프롬프트, 점수표, 원본 응답, 검증 도구까지 모두 공개해 결과를 그대로 재현할 수 있게 했다.

METAL LAB 해설 도표

FlavourBench 평가 구조

증거 상태측정 결과가 보고됨

  1. 문제 생성8개 재료 중 3개를 고르는 문제를 만들고 Epicure가 56가지 조합 전부에 미리 0~100점 점수를 매겨 얼린다
  2. 27개 모델 실행모든 모델이 동일한 534개 문제에 답하고, 완료된 문제만 공통 코어로 채택해 14,418개 응답 칸을 구성한다
  3. 채점모델 답변의 3개 재료 조합을 얼려둔 점수표에서 그대로 조회해 점수를 매긴다, 별도 채점 AI나 사람 평가 없음
  4. 통계 검증5만 번 부트스트랩으로 신뢰구간을, 10만 번 재표집과 Holm 보정으로 351개 모델 짝 비교를 계산한다
  5. 공개 검증 도구프롬프트, 점수표, 원본 응답, 경로 기록을 모두 공개해 오프라인에서 모든 결과를 재구성할 수 있게 한다
METAL LAB이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 사람 평가단이나 다른 AI 모델이 채점하던 기존 방식 대신, Epicure라는 요리 도메인 시스템이 8개 재료 중 3개를 고르는 모든 경우의 수(56가지)에 미리 점수를 매겨두고 그 표를 그대로 조회해 채점한다.
  2. 대체 재료 찾기, 재료 짝짓기, 제약이 있는 구성 등 세 가지 문제 유형에서 총 534개 문제를 만들고, GPT-5.6 계열, Claude, Gemini, Grok, Llama, Qwen, DeepSeek 등 27개 모델에게 똑같이 풀게 했다.
  3. 모든 모델이 모든 문제를 완료한 경우만 비교에 넣어(모델마다 534개씩, 총 14,418칸) 일부 모델만 쉬운 문제를 골라 답했다는 의심을 없앴다.
  4. 점수의 신뢰구간은 5만 번, 모델 간 우열 비교는 10만 번의 통계적 재표집(부트스트랩)으로 계산했고, 여러 번 비교할 때 생기는 우연한 오류를 막는 Holm 보정을 적용했다.
  5. 서로 독립적으로 만든 두 문제 세트로 같은 순위가 재현되는지도 확인했다.
Figure 2: FlavourBench Score with simultaneous 95% max-t intervals. The letter-free group numbers are inferential tiers: a point rank is shown for navigation, while models inside an unresolved group should not be read as a statistically established ordering.
Figure 2: FlavourBench Score with simultaneous 95% max-t intervals. The letter-free group numbers are inferential tiers: a point rank is shown for navigation, while models inside an unresolved group should not be read as a statistically established ordering.
Table 1: The three complete-core task families share one response and scoring interface but probe different culinary decisions. Each task contains eight candidates and 56 frozen scores.
FamilyDecisionEpicure utilityHard constraintTasks
Substitutionchoose a three-item replacement portfolio for an anchor ingredient0.8 anchor similarity +0.2 portfolio coherencenone178
Pairingchoose three ingredients to accompany an anchor0.65 anchor affinity +0.35 portfolio coherencenone178
Constraintschoose a feasible portfolio under diet and processing limits0.7 anchor similarity +0.3 coherencediet and maximum NOVA level178
Figure 3: Panel-1 versus panel-2 FlavourBench Score for every endpoint. The diagonal marks exact agreement. Each point aggregates 89 tasks per family; the joint analysis uses 534 unique anchor clusters.
Figure 3: Panel-1 versus panel-2 FlavourBench Score for every endpoint. The diagonal marks exact agreement. Each point aggregates 89 tasks per family; the joint analysis uses 534 unique anchor clusters.
Table 2: Frozen task-map diagnostics for 178 scheduled tasks per family, computed before model execution. Chance is the uniformly random portfolio mean; top gap is the median best–second-best margin.
FamilyExact chanceTop gapDistinct scores
Substitution45.25.656
Pairing45.05.656
Constraints3.437.34
Figure 4: All paired model contrasts after Holm correction. Teal means the row model scores higher; red means lower; grey is unresolved.
Figure 4: All paired model contrasts after Holm correction. Teal means the row model scores higher; red means lower; grey is unresolved.
Table 3: The complete automated leaderboard. Every score and every pairwise comparison uses the identical 534-task core. Rank intervals come from the anchor-cluster bootstrap; groups summarize the Holm-controlled pairwise graph.
RankModelFB Scoresimultaneous 95% CIrank 95% CIgroup
1Grok 4.665.1[61.0, 69.2][1, 5]1
2Gemini 3.1 Pro65.0[60.8, 69.1][1, 6]1
3GPT-5.6 Sol Pro64.2[60.1, 68.4][1, 8]1
4Muse Spark 1.263.8[59.6, 67.9][1, 10]1
5GPT-5.6 Terra Pro63.7[59.5, 67.8][1, 11]1
6Claude Fable 563.4[59.2, 67.5][2, 13]1
7GPT-5.6 Luna Pro62.6[58.5, 66.8][4, 14]1
8Claude Opus 562.5[58.4, 66.6][3, 15]1
9Qwen3.8 A95B62.1[57.9, 66.2][4, 17]1
10Kimi K362.1[57.9, 66.2][5, 16]1
11Gemini 3.6 Flash62.0[57.7, 66.2][4, 17]1
12DeepSeek V4 Pro 081362.0[57.8, 66.1][5, 17]1
13Qwen 3.8 Max61.5[57.4, 65.6][7, 18]1
14Tencent HY 361.5[57.3, 65.7][6, 19]1
15MiniMax M360.9[56.8, 65.1][8, 20]1
16GLM 5.360.6[56.4, 64.8][8, 20]1
17Muse Glimmer 30B59.9[55.9, 63.9][12, 21]2
18Seed 2.1 Turbo59.7[55.4, 64.0][12, 22]2
19Inkling59.6[55.4, 63.8][12, 22]2
20Claude Sonnet 559.5[55.4, 63.7][13, 22]2
21GLM 5.258.5[54.2, 62.7][16, 23]2
22Nemotron 3.5 Lightning57.4[53.2, 61.6][19, 25]2
23Command A56.7[52.5, 60.9][20, 25]2
24DeepSeek V4 Flash55.4[51.1, 59.7][22, 26]2
25Mistral Large 355.4[51.4, 59.5][22, 26]2
26Llama 4 Maverick53.7[49.6, 57.7][24, 26]3
27Command R+47.9[43.7, 52.0][27, 27]3
Figure 5: Mean score by family on the common core. Each family contributes 178 tasks per model. Values share the 0–100 scale but retain family-specific decision semantics.
Figure 5: Mean score by family on the common core. Each family contributes 178 tasks per model. Values share the 0–100 scale but retain family-specific decision semantics.
Table 4: Three common-core response examples. Parentheses give the fixed task score. The complete prompts and unabridged responses are released in machine-readable form.
FamilyPromptHigher-ranked selectionLower-ranked selection
SubstitutionSelect three alternatives to ‘boursin cheese‘ that best preserve its dairy role, regional context, and mutual portfolio coherence.caciocavallo, fromage blanc, grana padano (100)fromage blanc, quail egg, goose egg (15)
PairingSelect the three-ingredient bundle with the strongest learned pairing to ‘sweetbread‘ and the best internal coherence.parsnip, cipollini onion, porcini mushroom (92)parsnip, pickled onion, peppadew pepper (0)
ConstraintsFor a dish centered on ‘potato‘, select three vegetarian complements with NOVA processing level at most 2; among valid portfolios, maximize learned pairing and coherence.parsley, black pepper, thyme (100)cheese, bay leaf, parsley (0)
Figure 6: Point rank and 95% bootstrap rank interval for every endpoint. Narrow intervals imply a stable location; overlapping intervals mark unresolved local order.
Figure 6: Point rank and 95% bootstrap rank interval for every endpoint. Narrow intervals imply a stable location; overlapping intervals mark unresolved local order.

실제로 확인된 결과

  • 27개 모델을 동일한 534개 문제로 평가한 결과 Grok 4.6이 65.1점으로 가장 높았고, 동시 95% 신뢰구간은 61.0~69.2였다.
  • 독립적으로 만든 두 문제 세트 간 모델 점수 상관은 0.89, 순위 상관은 0.80으로 나타났다.
  • 351개 모델 짝 비교 중 Holm 보정 후 통계적으로 우열이 확정된 경우는 101개였다.
  • 전체 모델이 완료한 534개 문제만 리더보드에 포함시켜 모델마다 534칸씩, 총 14,418개 응답 칸을 채웠다.

어디에 쓸 수 있나

  • 언어모델의 개방형 의사결정 능력을 사람 평가나 다른 모델 채점 없이 자동으로, 재현 가능하게 비교하는 리더보드 설계에 참고할 수 있다.
  • 공개된 프롬프트·점수표·검증 도구를 활용해 다른 연구자가 결과를 재현하거나 새 모델을 같은 기준으로 추가 평가할 수 있다.
  • 56가지 조합에 대한 조밀한 보상 점수를 지도학습, 선호 학습, 강화학습 등 모델 훈련용 신호로 활용하는 다음 실험을 설계할 수 있다.

한계와 남은 검증

  • 점수는 하나의 Epicure 버전이 정의한 기준과의 일치도이며, 보편적인 사람의 미식 취향을 의미하지 않는다.
  • 평가 대상은 제한된 재료 선택 문제이며 전체 레시피 생성, 실제 조리, 안전 조언, 장기 주방 계획은 다루지 않는다.
  • 모든 모델이 완료한 문제만 채택하는 구조라 평가 범위가 그 공통 문제 집합으로 제한되며, 나머지는 별도 보충 트랙으로만 공개된다.
  • 결과는 특정 시점에 특정 경로(route)로 호출된 모델 버전에 묶여 있어 이후 모델 업데이트나 다른 호출 경로에는 그대로 적용되지 않는다.
  • Epicure 보상을 직접 최적화하는 학습이 실제 사람의 요리 결과로 이어지는지는 이 논문에서 검증하지 않았다.

왜 중요한가

모델 답변을 다른 AI나 사람이 평가하면 평가자 자체가 편향되거나 테스트 대상과 얽힐 수 있는데, 이 벤치마크는 미리 정해진 점수표로 자동 채점해 그런 문제를 줄인다. 순위 하나만 보여주지 않고 통계적으로 확실한 차이와 확실하지 않은 차이를 구분해서 보여주기 때문에, 리더보드 숫자를 그대로 믿기보다 신중하게 읽어야 한다는 점도 함께 알려준다.

이 논문의 용어

  • Epicure · 1,790개 재료를 300차원 공간으로 표현해 대체·짝짓기·식이 적합성 등을 계산하는 요리 시스템으로, 이 벤치마크의 정답 채점 기준이 된다
  • 앵커 클러스터 부트스트랩 · 같은 재료를 공유하는 문제들을 하나로 묶어 반복 재표집하며 신뢰구간을 계산하는 통계 방법
  • Holm 보정 · 여러 번 비교할 때 우연히 유의미해 보이는 결과가 늘어나는 것을 막기 위해 기준을 엄격하게 조정하는 통계 절차
  • 동시 95% 신뢰구간 · 27개 모델 점수 전체를 한꺼번에 봐도 95% 확률로 맞는다고 보장되는 구간
  • 제약이 있는 구성(constrained composition) · 식이 제한 같은 조건을 만족해야 점수를 받는 문제 유형으로, 조건을 못 맞추면 0점 처리된다

저자 · Josef Chen, Erim Hayretci

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Josef Chen et al., arXiv:2608.20574, CC BY 4.0