AI 과학 에이전트 9종을 실제 사용자 요청 178건으로 테스트했더니, 1등도 '합격선'을 확실히 넘지 못했다
arXiv:2608.216012026-08-25
K-Bench: measuring model performance on real scientific agent requests
AI 과학 에이전트 9종을 실제 사용자 요청 178건으로 테스트했더니, 1등도 '합격선'을 확실히 넘지 못했다
K-Bench 01은 시험문제나 정답이 정해진 과제 대신, 실제 서비스에서 사용자가 보낸 첨부파일 딸린 첫 메시지 178건을 그대로 가져와 아홉 개 최신 AI 모델에게 똑같은 환경에서 풀게 했다. 세 개의 AI 판정단이 8개 항목 기준으로 1,602건의 실행 결과를 채점한 결과, '전문가가 약간만 고치면 받아들일 수준'이라는 8점 기준선을 세 판정단 모두에게서 넘긴 모델은 하나도 없었다. 특히 결과물을 실제로 얼마나 잘 만들었는지보다 얼마나 말을 잘했는지에서 점수가 더 잘 나오는, 즉 '실제보다 잘한 것처럼 보이는' 경향이 아홉 모델 전부에서 공통으로 나타났다.
METAL LAB 해설 도표
K-Bench 01 평가 파이프라인
증거 상태측정 결과가 보고됨
실제 요청 수집K-Dense Web 실사용 트래픽에서 첨부파일이 딸린 사용자 첫 메시지 178건을 정답 없이 그대로 추출
동일 환경에서 실행아홉 개 최신 AI 모델이 동일한 샌드박스와 도구 세트로 178개 과제를 각각 수행, 총 1,602건 실행 완료
AI 판정단 채점신원을 가린 세 개의 AI 판정단이 실행 기록과 결과 파일을 열어보고 8개 항목 루브릭으로 39,934건 채점
결과 분석최고 점수 모델도 8점 기준선을 확실히 넘지 못했고, 과장(overclaiming)이 가장 흔한 실패 유형으로 확인
METAL LAB이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
K-Dense Web이라는 실제 서비스에 사용자가 보낸 요청 중 178건을 첨부파일까지 그대로 뽑아, 정답 없이 아홉 개 최신 AI 모델(gpt-5.6-sol, claude-opus-5 등)에게 동일한 샌드박스 환경에서 끝까지 수행시켰다.
완료된 1,602건의 실행 결과를 신원을 가린 세 개의 AI 판정단이 과제 완수도, 과학적 정확성, 도구 사용, 데이터 처리, 결과물 품질, 소통, 정직성 등 8개 항목과 종합 점수로 채점했고, 총 39,934건의 채점 값이 나왔다.
가장 높은 점수를 받은 gpt-5.6-sol은 평균 8.04점(10점 만점)이었지만 95% 신뢰구간이 7.80~8.23으로 '합격선' 8점을 확실히 넘지 못했고, 판정단 세 곳 중 두 곳은 claude-opus-5를 1위로 평가해 순위 자체가 판정단에 따라 흔들렸다.
전체 채점값의 47.6%가 8점 미달이었고, 아홉 모델 모두에서 예외 없이 과학적 정확성(평균 6.22점)이 소통 능력(평균 7.33점)보다 낮게 나와 '내용보다 표현이 앞서는' 경향이 일관되게 확인됐다.
가장 흔한 실패 유형은 실제로 한 것보다 부풀려 주장하는 '과장(overclaiming)'으로 전체 평가의 31.4%에서 나타났으며, 결과 파일을 아예 남기지 않은 실행도 상당수였다.
Figure 1: Graphical abstract. K-Bench 01 takes the first message of 178 real user sessions, verbatim and with attachments, runs each one end to end under nine frontier models in identical sandboxes, and has three identity-blinded judges score the resulting 1,602 runs after opening the artifacts each run left behind. The file icons in the left panel are illustrative of the attachment mix rather than a per-domain format breakdown; the most frequently attached format across the corpus is .docx (Table 30). Tile values are the headline results: the best model mean (8.04/10, a point estimate whose interval spans the acceptable line, and which only one of the three judges produces; see Section 5.4, and Section 5.2 for why the count of models reaching 8 is zero, one or two depending on the judge), the share of the 39,934 scored judgments below the acceptable line (47.6%, rendered as 48%), the share of assessments carrying the overclaiming tag (31.4%), and the share of tasks no model solved (12.4%, rendered as 12%). The strip beneath contrasts the mean of the execution dimensions (6.93) with the mean of the substance dimensions (6.34) and Section 4.2 gives the sharper and denominator-matched version, namely that scientific accuracy trails communication by 1.11 points within every one of the nine models.
Table 2: Self-presentation versus substance, by model. All values are model means over 4,806 assessments. The last two columns are the differences honesty − accuracy and communication − artifact quality; positive values mean the run reads better than it is. Computed from scores_wide.csv.
Model
Honesty
Sci. accuracy
Communication
Artifact quality
Hon.−Acc.
Comm.−Art.
nemotron-3-ultra-550b-a55b
6.29
3.62
3.85
1.38
+2.67
+2.48
gemma-4-31b-it
5.96
4.69
6.54
2.09
+1.26
+4.45
muse-spark-1.2
6.48
5.22
5.43
3.07
+1.26
+2.36
gpt-5.6-luna
8.78
7.64
8.51
6.76
+1.13
+1.75
gpt-5.6-sol
9.10
8.12
8.80
7.81
+0.98
+0.98
grok-4.5
7.79
6.85
8.31
6.78
+0.93
+1.53
kimi-k3
7.88
6.98
8.35
6.95
+0.90
+1.40
claude-opus-5
8.00
7.31
8.38
7.62
+0.69
+0.77
gemini-3.6-flash
5.41
5.49
7.76
5.70
−0.08
+2.06
Figure 2: Representative benchmarks for scientific and agentic capability, 2020–2026, grouped by what they measure. Horizontal positions are the first public posting dates of the cited work; the figure is a positioning aid, not an exhaustive census, and several suites could reasonably sit in more than one lane. Several suites discussed below are omitted here for legibility; Table 1 gives the fuller comparison. They include BenchBench-Protocol, which is the closest relative of K-Bench in construction philosophy and would sit in the middle lane at 2026. K-Bench 01 is placed at the right. Building a benchmark out of deployment traffic is not itself new: WildBench and Arena-Hard curate items from chat logs and RealClawBench reconstructs developer-agent sessions (Lin et al. 2024; Li et al. 2024; Lv et al. 2026), and within science AstaBench is the nearest precedent, with problems inspired by requests to its deployed agents. What distinguishes K-Bench 01 is that its items are users’ first turns verbatim, with the attachments and without reconstruction, and that grading opens the files a run produced rather than matching a reference answer.
Table 3: Failure-mode taxonomy from rubric v1.0. Judges apply every tag that fits and may apply none.
Tag
Definition
premature_completion
Stopped and declared done while major work remained.
fabricated_results
Presented numbers/findings that were never computed (placeholders, invented values, simulated data passed off as real).
fabricated_citations
Cited papers, datasets, or sources that don’t exist or don’t support the claim.
ignored_attachments
User-provided files were not opened or not used when the task required them.
misread_data
Files were opened but parsed or interpreted incorrectly (wrong columns, wrong units, wrong sheet).
tool_thrashing
Extended loops of near-identical failing commands with no strategy change.
environment_failure_unrecovered
A missing package/dependency/resource blocked progress and the agent never found a workaround.
wrong_language
Response not in the language of the user’s prompt (substantially).
truncated_run
The run was cut off before the agent finished (use with the truncation banner).
scope_drift
Did substantial work the user didn’t ask for while neglecting what they did ask for.
missing_artifacts
Promised or clearly-required output files were not produced.
Superficial treatment where the task demanded depth (e.g., generic textbook answer to a specific data question).
overclaiming
Final answer overstates quality, completeness, or certainty of what was done.
format_noncompliance
Ignored an explicit format request (file type, structure, template, length).
other
Anything else — must be explained in the summary.
Figure 3: Mean overall score by model, pooled over judges. Blue bars are means over the 178 tasks with a 95% bootstrap interval marked by the white tick; the gray bar spans the strictest to the most lenient judge’s mean for that model. The shaded region marks scores at or above the rubric’s 8-anchor. The panel title printed inside the figure states that one model reaches that line; that count is the pooled-panel value, and it is zero, one or two depending on which judge is asked (Section 5.2).
Table 4: Headline results by model. Overall is the mean across 178 tasks of the three-judge mean, with a 95% percentile bootstrap interval over sessions. The three judge columns give the same quantity computed from that judge alone. Majority success requires more than half of the three judges to independently mark the run fully successful; unanimous requires all three. “Scores ≥8” is the share of that model’s individual scored judgments at or above the acceptable line. Computed from scores_wide.csv.
Model
Overall (95% CI)
gpt-5.6-sol
qwen3.8-max
grok-4.5
Majority
Unanimous
Scores ≥8
gpt-5.6-sol
8.04 [7.80, 8.23]
7.90
8.20
8.01
71%
56%
89%
claude-opus-5
7.61 [7.40, 7.82]
6.40
8.29
8.15
75%
23%
79%
gpt-5.6-luna
7.46 [7.23, 7.68]
7.17
7.71
7.49
60%
42%
77%
kimi-k3
7.17 [6.95, 7.38]
6.15
7.81
7.54
52%
15%
69%
grok-4.5
6.96 [6.74, 7.17]
6.15
7.46
7.26
46%
20%
64%
gemini-3.6-flash
5.84 [5.60, 6.06]
4.80
6.59
6.13
25%
7%
36%
muse-spark-1.2
4.27 [3.91, 4.65]
3.76
4.52
4.54
17%
4%
25%
gemma-4-31b-it
3.86 [3.57, 4.15]
3.61
3.94
4.02
6%
2%
16%
nemotron-3-ultra-550b-a55b
2.78 [2.38, 3.14]
2.42
2.93
3.00
9%
3%
14%
Figure 4: The same 1,602 runs scored by three judges on three different scales. Each panel gives mean overall score by model for one judge, with 95% bootstrap intervals over sessions; rows are held in the order of the pooled mean.
Table 5: Score distribution by judge, counts over the eight rubric dimensions plus the holistic overall, N/A excluded. Computed from scores_wide.csv.
Judge
n
0
1
2
3
4
5
6
7
8
9
10
≥8
<5
gpt-5.6-sol
13,271
631
400
500
792
1,117
1,179
1,481
2,010
2,908
2,014
239
38.9%
25.9%
grok-4.5
13,240
424
410
448
548
611
855
859
1,527
4,109
3,215
234
57.1%
18.4%
qwen3.8-max
13,423
432
442
345
468
495
818
947
1,266
3,434
4,531
245
61.2%
16.3%
Figure 5: How often a run fully satisfies the scientist who asked. Light bars give the share of that model’s runs marked fully successful by a majority of the three judges; dark bars require unanimity; the gray whisker spans the individual judges. claude-opus-5 has the highest majority rate of any model (75%) while gpt-5.6-sol has 71%. The two differ far more on the unanimous rate (23% against 56%), but that gap should not be read as a property of the models: unanimity requires the strictest judge to agree, that judge is gpt-5.6-sol, and it scores claude-opus-5 1.8 points below its peers (Sections 5.3 and 5.4).
Table 6: Rubric dimensions, hardest to easiest. Mean is pooled over all models, judges and tasks. Percentages use non-N/A denominators, so n is the applicable count for that dimension and the two conditional dimensions are scored only where they applied. Per-model columns use abbreviated names, left to right in leaderboard order. Computed from scores_wide.csv.
Dimension
Mean
≥8
<5
n
sol
opus
luna
kimi
grok
gemini
muse
gemma
nemotron
Artifact quality
5.50
40.6%
32.4%
3,094
7.81
7.62
6.76
6.95
6.78
5.70
3.07
2.09
1.38
Scientific accuracy
6.22
42.7%
23.5%
4,806
8.12
7.31
7.64
6.98
6.85
5.49
5.22
4.69
3.62
Task fulfillment
6.41
51.2%
25.2%
4,806
8.23
8.12
7.78
7.62
7.45
6.77
4.64
4.12
2.94
Reasoning quality
6.68
48.8%
19.5%
4,806
8.48
8.45
7.97
7.65
7.38
6.27
5.28
4.29
4.39
Tool use
6.79
53.8%
16.2%
4,806
8.07
8.46
7.61
7.84
7.75
6.90
5.00
4.32
5.13
Data handling
7.14
59.8%
13.5%
3,198
8.66
8.47
8.18
8.02
7.57
6.80
6.36
4.17
5.04
Honesty / calibration
7.30
59.7%
13.1%
4,806
9.10
8.00
8.78
7.88
7.79
5.41
6.48
5.96
6.29
Communication
7.33
70.7%
11.9%
4,806
8.80
8.38
8.51
8.35
8.31
7.76
5.43
6.54
3.85
Figure 6: Distribution of individual scored judgments by judge. Every dimension score and the holistic score from every judged run, pooled per judge; the axis labels shorten this to “dimension” scores. The strictest judge places 61.1% of scores below the acceptable line and the most lenient 38.8%, so even the lenient reading leaves well over a third of the work short of acceptable.
Table 7: Failure-mode frequencies, as a percentage of judged runs (assessments). Tags are not exclusive. Computed from the failure_modes field of scores_wide.csv.
Failure mode
All
gpt-5.6-sol
claude-opus-5
gpt-5.6-luna
kimi-k3
grok-4.5
gemini-3.6-flash
muse-spark-1.2
gemma-4-31b
nemotron-3
overclaiming
31.4
6.4
32.6
9.7
31.8
27.5
68.2
34.6
44.2
27.3
missing_artifacts
22.6
6.4
12.4
9.7
13.5
11.0
12.2
39.5
40.3
58.1
shallow_analysis
17.7
2.6
0.9
5.6
8.4
11.0
30.0
16.3
60.5
23.8
truncated_run
16.4
0.0
4.5
2.8
2.8
11.2
1.1
52.8
7.9
64.6
premature_completion
12.5
4.9
3.6
7.7
5.2
6.0
8.4
15.2
40.4
21.5
statistical_malpractice
6.6
1.5
7.9
1.9
6.9
9.2
16.3
6.4
7.1
2.1
fabricated_results
5.5
0.4
3.0
0.4
2.2
4.7
21.3
5.8
9.6
2.2
tool_thrashing
5.1
0.6
0.0
0.6
0.0
0.0
2.8
30.5
0.2
11.4
format_noncompliance
5.0
2.1
4.1
3.9
3.9
3.2
7.7
5.8
10.9
3.7
other
4.4
1.9
4.7
5.2
4.7
3.4
6.7
2.8
5.2
4.7
fabricated_citations
2.9
0.4
0.9
0.6
2.2
2.4
11.2
2.8
2.8
2.8
ignored_attachments
2.7
0.4
0.0
0.7
0.2
0.6
3.0
2.2
12.7
4.7
environment_failure_unrecovered
2.2
1.5
1.3
2.1
1.5
0.9
1.1
2.1
4.5
4.5
misread_data
2.1
0.9
1.1
0.6
2.2
1.5
5.6
0.7
3.6
2.2
wrong_language
1.1
0.0
1.1
0.2
0.0
0.0
0.7
0.4
1.7
6.0
scope_drift
0.4
0.6
0.2
0.6
0.0
0.0
0.7
0.0
0.6
1.1
Figure 7: Mean score by rubric dimension across every model, judge and task. Execution dimensions (blue) average 6.93; substance dimensions (orange) average 6.34. Task fulfillment and data handling (gray) belong cleanly to neither group. The color assignment is a judgment call and the 0.59-point gap is sensitive to it (Section 4.2); the denominator-matched comparison of scientific accuracy against communication is not. No dimension reaches the acceptable line on average.
Table 8: Share of runs that used each tool at least once (%). Computed from the tool_calls_by_tool field of run_metrics.csv over all 1,602 runs.
Tool
All
sol
opus
luna
kimi
grok
gemini
muse
gemma
nemotron
bash
75
87
96
79
75
83
83
68
43
66
read
45
69
66
61
52
35
30
34
30
31
write
43
51
75
42
61
46
29
35
23
28
web_search
38
66
60
53
38
30
25
28
16
26
fetch_content
26
66
40
49
23
20
2
13
4
12
edit
24
44
48
40
33
21
11
2
13
3
get_search_content
18
53
31
33
11
16
1
9
3
6
source_check
9
34
30
10
10
1
0
1
0
0
Figure 8: Every model against every dimension. Columns run weakest to strongest, rows strongest to weakest. Artifact quality is the weakest dimension for seven of the nine models; the exceptions are claude-opus-5, weakest on scientific accuracy (7.31 against 7.62), and gemini-3.6-flash, weakest on honesty and calibration (5.41 against 5.70). No single dimension is the strongest for a majority: communication tops four models, honesty and calibration four more, and data handling one.
Table 9: Deliverables on disk versus outcome. Left: runs grouped by whether any output file exists. Right: runs grouped by file count. Computed by joining run_metrics.csv to run-level means from scores_wide.csv.
Files?
n runs
Overall
Majority
File count
n runs
Overall
Majority
No
767
5.26
37.5%
0
767
5.26
37.5%
Yes
835
6.68
42.4%
1–2
259
6.87
45.2%
3–10
202
6.49
40.1%
11+
374
6.64
41.7%
Figure 9: Failure-mode frequency overall (left) and by model (right). The leading tag across the corpus is overclaiming at 31.4% of assessments. The per-model panel shows that this is not a property of the weakest systems only: gemini-3.6-flash carries it on 68% of its assessments and gemma-4-31b-it on 44%, but so do claude-opus-5 (33%) and kimi-k3 (32%), while gpt-5.6-sol (6%) and gpt-5.6-luna (10%) are markedly cleaner.
Table 10: Paired win rate of the row model against the column model (%), same task and same judge, ties excluded. 534 paired comparisons per cell. Computed from scores_wide.csv.
gpt-5.6-sol
claude-opus-5
gpt-5.6-luna
kimi-k3
grok-4.5
gemini-3.6-flash
muse-spark-1.2
gemma-4-31b
nemotron-3
gpt-5.6-sol
—
64
84
85
91
94
96
97
98
claude-opus-5
36
—
60
75
77
92
94
96
98
gpt-5.6-luna
16
40
—
63
74
89
93
96
98
kimi-k3
15
25
37
—
60
89
91
94
97
grok-4.5
9
23
26
40
—
86
93
94
97
gemini-3.6-flash
6
8
11
11
14
—
74
90
92
muse-spark-1.2
4
6
7
9
7
26
—
57
75
gemma-4-31b-it
3
4
4
6
6
10
43
—
70
nemotron-3-ultra-550b-a55b
2
2
2
3
3
8
25
30
—
Figure 10: Tool reach (left) and empty-handed runs (right). Left: the share of runs in which each model used each tool at least once, rows ordered by overall usage rather than by category; the evidence-seeking tools (web_search, fetch_content, get_search_content and source_check) separate the field most sharply. Right: the share of runs that end with no file on disk, from 17% for claude-opus-5 to 75% for nemotron-3-ultra-550b-a55b. A run can score well on prose and still leave a user with nothing.
Table 11: Bradley-Terry strengths fitted from the paired outcomes (log scale, centered, 95% bootstrap CI over sessions), and mean paired score delta against the leader with win and loss shares. Strengths are identified up to an additive constant, so only differences are interpretable.
Bradley-Terry
Versus gpt-5.6-sol
Model
Strength
CI low
CI high
Mean Δ
CI
Wins
Losses
gpt-5.6-sol
+1.73
+1.54
+1.90
—
—
—
—
claude-opus-5
+1.31
+1.17
+1.45
−0.42
[−0.64, −0.20]
0.22
0.39
gpt-5.6-luna
+1.02
+0.87
+1.16
−0.58
[−0.78, −0.38]
0.10
0.50
kimi-k3
+0.71
+0.58
+0.86
−0.87
[−1.07, −0.66]
0.11
0.61
grok-4.5
+0.54
+0.42
+0.65
−1.08
[−1.30, −0.87]
0.06
0.67
gemini-3.6-flash
−0.44
−0.56
−0.32
−2.20
[−2.43, −1.96]
0.05
0.83
muse-spark-1.2
−1.19
−1.36
−1.00
−3.76
[−4.15, −3.37]
0.04
0.89
gemma-4-31b-it
−1.56
−1.73
−1.38
−4.18
[−4.49, −3.85]
0.03
0.93
nemotron-3-ultra-550b-a55b
−2.12
−2.36
−1.90
−5.25
[−5.65, −4.85]
0.02
0.94
Figure 11: Measured behavior from the transcripts, no judge involved. Verification actions per run, code written per run, and distinct scientific packages imported per run. claude-opus-5 verifies its own output about 41 times as often per run as gemma-4-31b-it and writes roughly 23 times more code. Verification frequency is deterministic to compute and tracks the score ordering closely. Extracted from the 1,602 run transcripts.
Table 12: Paired win rate against gpt-5.6-sol by rubric dimension (%), same task and same judge, ties excluded. For the two conditional dimensions, pairs in which either run was marked not-applicable are dropped, so those rows rest on fewer decisive pairs than the other six. Computed from scores_wide.csv.
Dimension
opus
luna
kimi
grok
gemini
muse
gemma
nemotron
Tool use
73.2
26.2
38.4
28.1
12.9
8.2
3.3
4.5
Reasoning quality
49.1
17.1
13.7
6.7
4.7
3.4
2.3
2.0
Task fulfillment
46.1
18.8
18.6
15.0
11.4
6.4
2.8
2.0
Data handling
44.6
12.1
11.9
9.0
5.8
2.1
0.8
0.6
Artifact quality
40.3
13.2
17.6
12.3
8.2
2.5
1.6
1.3
Communication
30.0
17.1
20.6
13.2
5.3
2.9
2.1
2.5
Scientific accuracy
22.5
18.6
11.1
5.9
3.0
2.1
1.8
1.6
Honesty / calibration
17.2
22.2
7.5
6.8
0.6
1.9
4.6
2.3
Figure 12: Paired comparison matrix (left) and fitted Bradley-Terry strengths with 95% bootstrap intervals (right). Each cell of the matrix aggregates 534 comparisons on the same task by the same judge. Pairing removes the additive judge-calibration term that dominates the pooled means and separates adjacent models that the means leave overlapping.
Table 13: Mean overall by domain and model. Computed from scores_wide.csv joined to the domain field of run_metrics.csv.
Domain
n tasks
sol
opus
luna
kimi
grok
gemini
muse
gemma
nemotron
Chemistry, drug, materials
17
8.33
7.75
7.69
7.02
6.94
5.57
4.00
3.18
2.75
Clinical and health
59
7.80
7.61
7.57
7.20
7.14
6.08
4.40
3.95
3.33
Life sciences
59
8.21
7.81
7.42
7.14
6.79
5.52
4.14
3.56
2.72
Physical sciences, eng., CS
43
8.00
7.30
7.26
7.22
6.94
6.05
4.39
4.40
2.13
Table 14: Effect of attachments on mean overall, by model. 125 of 178 sessions carry at least one file. Computed by joining run_metrics.csv to run-level means.
Model
No attachments
Has attachments
Δ
gemma-4-31b-it
4.93
3.40
−1.52
muse-spark-1.2
5.29
3.84
−1.45
nemotron-3-ultra-550b-a55b
3.69
2.40
−1.29
gemini-3.6-flash
6.23
5.68
−0.56
claude-opus-5
7.73
7.57
−0.16
gpt-5.6-luna
7.52
7.43
−0.09
grok-4.5
7.00
6.94
−0.06
gpt-5.6-sol
7.98
8.06
+0.08
kimi-k3
6.84
7.30
+0.46
Table 15: Mean overall by prompt-length quartile. Quartiles are formed over the 178 tasks by the byte length of the first user message, which is held in the task registry rather than in the score tables (Section Data, code and availability). Eight of the nine models are tabulated here for width; gemma-4-31b-it is plotted alongside them in Figure 13 and also scores lower on Q4 than on Q1.
Bin
Median bytes
n
sol
opus
luna
kimi
grok
gemini
muse
nemotron
Q1
96
45
8.31
7.96
7.76
7.18
7.49
6.38
5.47
4.87
Q2
284
44
8.23
7.55
7.52
7.24
7.04
5.91
4.25
2.70
Q3
1,232
45
8.09
7.70
7.55
7.23
6.79
5.82
3.66
2.02
Q4
6,217
44
7.50
7.23
6.99
7.02
6.49
5.24
3.69
1.51
Q1 → Q4 drop
0.8
0.7
0.8
0.2
1.0
1.1
1.8
3.4
Table 16: Run endings and outcomes. Left: mean overall score and majority-success rate by final stop reason. Right: truncation rate by model, with that model’s mean overall for reference. Computed from run_metrics.csv joined to run-level means.
Ending
n
Overall
Majority
Model
Truncated
Overall
stop
1,354
6.70
47.4%
nemotron-3-ultra-550b-a55b
64.6%
2.78
length
224
2.22
0.0%
muse-spark-1.2
52.8%
4.27
toolUse
10
3.60
0.0%
grok-4.5
11.2%
6.96
error
14
0.00
0.0%
gemma-4-31b-it
7.9%
3.86
Not truncated
1,344
6.73
47.8%
claude-opus-5
3.9%
7.61
Truncated
258
2.18
0.0%
kimi-k3
2.8%
7.17
gpt-5.6-luna
1.7%
7.46
gpt-5.6-sol
0.0%
8.04
gemini-3.6-flash
0.0%
5.84
Table 17: Run-level Spearman correlates of overall score (n=1,602). Computed from run_metrics.csv joined to run-level means.
Variable
ρ
Variable
ρ
Input tokens
−0.42
Turns
+0.23
Thinking characters
+0.35
Output files
+0.21
Wall-clock seconds
+0.33
Tool error rate
−0.16
Cost (USD)
+0.31
Attachment count
−0.16
Tool calls
+0.27
Table 18: Cost and efficiency. Mean and median inference cost per task, total campaign cost, and two efficiency ratios. Computed from run_metrics.csv; total generation cost across all 1,602 runs was $3,649.18.
Model
Overall
Mean $
Median $
Total $
Majority
Score/$
Maj. pts/$
gpt-5.6-sol
8.04
8.51
4.72
1,514.92
71%
0.9
8.4
claude-opus-5
7.61
6.69
4.04
1,190.15
75%
1.1
11.3
gpt-5.6-luna
7.46
0.15
0.06
26.80
60%
49.5
395.6
kimi-k3
7.17
1.11
0.40
197.51
52%
6.5
46.6
grok-4.5
6.96
0.38
0.23
67.96
46%
18.2
119.2
gemini-3.6-flash
5.84
0.82
0.50
146.45
25%
7.1
30.0
muse-spark-1.2
4.27
2.42
0.27
430.57
17%
1.8
7.2
gemma-4-31b-it
3.86
0.02
0.00
3.68
6%
186.4
298.7
nemotron-3-ultra-550b-a55b
2.78
0.40
0.08
71.14
9%
7.0
22.5
Table 19: Failure co-occurrence, P(column∣row) in percent. Read a row as: when this failure occurs, how often the column failure also occurs. Computed from the failure_modes field of scores_wide.csv.
Given
overclaim.
missing_art.
shallow
truncated
premature
stat. malp.
overclaiming
100
20
37
10
15
19
missing_artifacts
28
100
23
48
38
8
shallow_analysis
65
29
100
15
33
10
truncated_run
19
65
16
100
20
2
premature_completion
38
68
46
26
100
7
statistical_malpractice
92
28
27
5
14
100
Table 20: Pairwise judge agreement on the holistic overall score (n=1,602 runs). A−B is the mean difference in level. Computed from scores_wide.csv.
Judge A
Judge B
Spearman ρ
Mean |A−B|
Within ±1
A−B
gpt-5.6-sol
qwen3.8-max
0.78
1.37
62.5%
−1.01
gpt-5.6-sol
grok-4.5
0.83
1.18
69.0%
−0.87
qwen3.8-max
grok-4.5
0.89
0.60
91.3%
+0.14
Table 21: Judge agreement by rubric dimension. Means over the three judge pairs, computed on runs where all three judges scored the dimension. Computed from scores_wide.csv; n is smaller for the two conditional dimensions because all three judges must have marked them applicable.
Dimension
n runs
Mean pairwise ρ
Mean |diff|
Within ±1
Artifact quality
971
0.86
0.92
77%
Task fulfillment
1,602
0.85
0.87
81%
Overall
1,602
0.83
1.05
74%
Reasoning quality
1,602
0.80
0.95
78%
Data handling
1,004
0.76
0.88
81%
Scientific accuracy
1,602
0.76
1.47
60%
Communication
1,602
0.76
0.59
91%
Tool use
1,602
0.74
0.94
81%
Honesty / calibration
1,602
0.69
1.32
65%
Table 22: Calibration-adjusted self-preference for the two judges that are also contestants. All values are mean overall scores. Computed from scores_wide.csv.
Judge
Scores itself
Scores others
Peers score it
Peers score others
Adjusted SP
gpt-5.6-sol
7.90
5.06
8.10
6.09
+0.83
grok-4.5
7.26
6.11
6.80
5.76
+0.11
Table 23: gpt-5.6-sol’s deviation from peer consensus, by model judged. “Peer mean” is the mean of the two other judges’ means for that model. A single additive strictness term would make the deviation column constant; it is not. The two smallest magnitudes belong to models the panel already places near 3, where downward deviation is bounded by the floor of the scale, which biases the pooled baseline — and hence the self-preference estimate of Table 22 — toward zero. Computed from the per-judge columns of Table 4.
Model judged
gpt-5.6-sol
Peer mean
Deviation
claude-opus-5
6.40
8.22
−1.82
gemini-3.6-flash
4.80
6.36
−1.56
kimi-k3
6.15
7.68
−1.53
grok-4.5
6.15
7.36
−1.21
muse-spark-1.2
3.76
4.53
−0.77
nemotron-3-ultra-550b-a55b
2.42
2.97
−0.55
gpt-5.6-luna (same family)
7.17
7.60
−0.43
gemma-4-31b-it
3.61
3.98
−0.37
gpt-5.6-sol (itself)
7.90
8.11
−0.21
Table 24: Exact systems evaluated. All models were accessed through OpenRouter; the identifier column is the slug passed to the API and the date is the model’s OpenRouter listing date, not a vendor snapshot date. Context length is the window advertised by the endpoint. All nine benchmarked models ran under the stock pi 0.84.0 harness in Modal sandboxes with the same tool set, with thinking level max requested where the model exposed one, no model-specific prompting, no sub-agents and no retries, against zero-retention endpoints (Section Data provenance, consent and privacy). The campaign ran between 6 and 12 August 2026. Context length does not predict truncation (Section 4.7).
Model as named here
Provider
OpenRouter identifier
Listed
Context
gpt-5.6-sol
OpenAI
openai/gpt-5.6-sol
9 Jul 2026
1,050,000
claude-opus-5
Anthropic
anthropic/claude-opus-5
24 Jul 2026
1,000,000
gpt-5.6-luna
OpenAI
openai/gpt-5.6-luna
9 Jul 2026
1,050,000
kimi-k3
Moonshot AI
moonshotai/kimi-k3
16 Jul 2026
1,048,576
grok-4.5
xAI
x-ai/grok-4.5
8 Jul 2026
500,000
gemini-3.6-flash
Google
google/gemini-3.6-flash
21 Jul 2026
1,048,576
muse-spark-1.2
Meta
meta/muse-spark-1.2
5 Aug 2026
1,048,576
gemma-4-31b-it
Google
google/gemma-4-31b-it
2 Apr 2026
262,144
nemotron-3-ultra-550b-a55b
NVIDIA
nvidia/nemotron-3-ultra-550b-a55b
4 Jun 2026
512,288
qwen3.8-max (judge only)
Alibaba
qwen/qwen3.8-max
3 Aug 2026
1,000,000
Table 25: Objective run metrics by model. Medians over 178 runs except where noted. “No output files” is the share of runs ending with an empty output tree. Computed from run_metrics.csv.
Model
Med. turns
Med. tool calls
Tool error rate
Med. minutes
Med. $/task
Total $
No output files
gpt-5.6-sol
40.5
59.5
0.04
17.9
4.72
1,514.92
37.1%
claude-opus-5
38.0
47.0
0.03
29.0
4.04
1,190.15
17.4%
gpt-5.6-luna
30.0
44.0
0.05
28.1
0.06
26.80
47.8%
kimi-k3
13.0
13.5
0.05
20.3
0.40
197.51
34.3%
grok-4.5
9.0
12.0
0.09
4.6
0.23
67.96
37.6%
gemini-3.6-flash
16.0
15.0
0.10
4.2
0.50
146.45
41.6%
muse-spark-1.2
11.0
11.0
0.26
2.1
0.27
430.57
67.4%
gemma-4-31b-it
2.0
1.0
0.06
2.1
0.00
3.68
73.0%
nemotron-3-ultra-550b-a55b
7.0
7.0
0.20
3.3
0.08
71.14
74.7%
실제로 확인된 결과
최고 점수 모델 gpt-5.6-sol의 평균은 8.04점(95% 구간 7.80~8.23)으로 세 판정단 중 어느 곳에서도 확실히 8점 기준선을 넘지 못했고, 두 판정단은 claude-opus-5를 1위로 평가했다.
전체 39,934건 채점값 중 47.6%가 8점 미달이었고, 20.2%는 5점 미달이었다.
판정단 majority 기준으로 완전히 성공했다고 평가된 실행은 40.1%, 세 판정단 전원 일치로는 19.2%였고, 45.9%는 전원 일치로 불만족 판정을 받았다.
178개 과제 중 22개(12.4%)는 아홉 모델 중 어느 것도 majority 성공 판정을 받지 못했다.
과학적 정확성 평균은 6.22점, 소통 평균은 7.33점으로 이 격차는 아홉 모델 모두에서 같은 방향으로 나타났으며, 가장 흔한 실패 태그는 과장(overclaiming, 31.4%)이었다.
어디에 쓸 수 있나
과학 연구 지원용 AI 에이전트를 도입하기 전, 순위표 1위 숫자만 보지 않고 실제로 파일·결과물을 만들어내는지, 주장이 과장되지 않았는지 함께 점검하는 평가 방식으로 참고할 수 있다.
회사나 연구실에서 자체 AI 도구를 평가할 때, 정답이 정해진 문제 대신 실제 사용자의 첫 요청(첨부파일 포함)을 그대로 사용하는 평가 설계 방식으로 참고할 수 있다.
AI 에이전트가 만든 보고서나 분석 결과를 검토할 때, 문장이 매끄럽다고 결과물의 정확성이나 완성도까지 좋다고 단정하지 않는 검증 습관을 세우는 데 참고할 수 있다.
한계와 남은 검증
이 평가는 K-Dense Web이라는 특정 서비스의 사용자 요청 178건에 한정되어 있어, 다른 종류의 과학 작업이나 다른 사용자층에는 그대로 적용되지 않을 수 있다.
정답이 없는 과제라 '정확도'는 절대적 진리가 아니라 채점 기준(루브릭)에 따른 상대적 판단이며, 전문가 인간 평가와의 비교 기준선도 없다.
각 실행은 한 번씩만 수행되고 채점되어, 모델의 실제 능력과 실행마다의 우연한 변동을 구분할 수 없다.
판정단이 AI 모델이기 때문에 판정단 자체의 편향(자기 선호, 문체 편향 등)이 결과에 영향을 줄 수 있으며, 판정단에 따라 순위와 절대 점수가 달라지는 문제가 완전히 해결되지 않았다.
과제 데이터셋 자체는 비공개로 유지되어, 다른 연구자가 동일한 과제로 재현하기는 어렵다.
왜 중요한가
지금까지 과학 AI 평가는 정답이 정해진 시험문제 위주였는데, 이 연구는 실제 사용자가 보낸 애매하고 파일 첨부된 요청으로 평가해 현실과 더 가까운 그림을 보여준다. AI 에이전트를 과학 업무에 실제로 도입하려는 사람들에게, 순위표 1위라는 숫자보다 '무엇을 실제로 만들어냈는지, 무엇을 과장했는지'를 함께 봐야 한다는 메시지를 준다.
이 논문의 용어
판정단(judge) · 정답을 대조하는 대신, 미리 정해진 채점 기준(루브릭)에 따라 결과물을 평가하는 또 다른 AI 모델
8-anchor(8점 기준선) · 채점 기준에서 '해당 분야 전문가가 약간만 고치면 받아들일 만한 수준'으로 정의된 점수 기준
overclaiming(과장) · 실제로 수행하거나 확인한 것보다 더 잘했다고 주장하는 실패 유형
페어드 비교(paired comparison) · 같은 과제, 같은 판정단이 채점한 두 모델의 결과를 짝지어 직접 비교해 판정단 차이와 과제 난이도 영향을 상쇄하는 방법