Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme
评分方式算错了,AI考试成绩就会虚高:用越南高考真实评分规则做的基准测试
越南2025年改革后的高中毕业考试中,有一部分采用非比例的阶梯式评分:四道判断题里答对三道只给0.50分,而不是常规正确率算法会给的0.75分。现有的AI基准测试大多忽略这种评分规则,直接用简单正确率打分,结果虚报了国家实际上不会认可的能力水平。研究团队用21场真实考试的632道题构建了THPT-Ladder基准,按官方评分规则测试了8个模型,发现正确率根本无法预测模型最终会得多少分。
他们做了什么
- 越南2025年改革后考试的第二部分要求考生判断每题四条判断陈述的对错,按答对的条数给出0、0.10、0.25、0.50、1.00分的阶梯式非比例评分,答对三条只得0.50分而非按比例应得的0.75分。
- 这部分占整场考试10分中的4分,如果像通常的基准测试那样按逐条陈述的正确率打分,得出的分数会一直高于教育部实际给出的分数。
- 研究团队构建了THPT-Ladder基准,包含11个学科21场官方考试中的632道题,完全按照教育部公布的标准答案和评分规则打分,并与超过一百万名真实考生的公开成绩分布对比,把模型得分换算成人群百分位。
- 对8个模型(3个开放权重模型、5个闭源模型)的测试显示,官方阶梯评分比按比例评分每题少给0.020到0.159分。以Qwen3.5-27B为例,在2025年历史科目考试中这一差距使其名次从第90百分位跌到第77百分位。
- 正确率相同的模型,若错误分布方式不同,最终得分可能差异很大;在与Claude Sonnet 5相同正确率水平下,每题得分可在0.869到0.932分之间浮动。
- 一个未读任何题目、随便选定的固定答案字符串(DDSS)利用官方答案分布的不均衡,就能拿到整场考试11.07%到24.25%的分数。
| structure | items | ||||||
|---|---|---|---|---|---|---|---|
| Subject | I | II | III | rnd | 2025 | 2026 | Fig. |
| Biology | 18 | 4 | 6 | 2.350 | 28 | 28 | 28 |
| Chemistry | 18 | 4 | 6 | 2.350 | 28 | 28 | 9 |
| Economics & Law | 24 | 4 | 0 | 2.725 | 28 | 28 | 0 |
| English | 40 | 0 | 0 | 2.500 | 40 | 40 | 0 |
| Geography | 18 | 4 | 6 | 2.350 | 28 | 28 | 8 |
| History | 24 | 4 | 0 | 2.725 | 28 | 28 | 1 |
| Informatics | 24 | 4a | 0 | 2.725 | 30 | 30 | 0 |
| Mathematics | 12 | 4 | 6 | 1.975 | 22 | 22 | 14 |
| Physics | 18 | 4 | 6 | 2.350 | 28 | 28 | 4 |
| Technology (Agri.) | 24 | 4 | 0 | 2.725 | 56 | 28 | 4 |
| Technology (Ind.) | 24 | 4 | 0 | 2.725 | 28 | n.p. | 12 |
| total | 632 items | 80 |
| Model | % avail. | I | II stmt | II pts/q | shf/q | III |
|---|---|---|---|---|---|---|
| Qwen3.5-27B | 93.8 | 0.973 | 0.926 | 0.884 | 0.042 | 0.917 |
| Qwen3.5-9B | 90.7 | 0.965 | 0.896 | 0.830 | 0.066 | 0.850 |
| InternVL3.5-8B | 69.3 | 0.838 | 0.756 | 0.597 | 0.159 | 0.183 |
| Claude Opus 5 | 97.0 | 0.988 | 0.973 | 0.949 | 0.024 | 0.950 |
| GPT-5.5 | 97.0 | 0.988 | 0.967 | 0.948 | 0.020 | 0.950 |
| Claude Sonnet 5 | 93.6 | 0.977 | 0.935 | 0.881 | 0.054 | 0.917 |
| Claude Opus 4.8 | 92.9 | 0.975 | 0.920 | 0.867 | 0.052 | 0.917 |
| Claude Haiku 4.5 | 82.9 | 0.934 | 0.842 | 0.735 | 0.108 | 0.567 |
为什么重要
如果把考试的真实评分规则简化成简单正确率,AI模型可能会显得通过了一场它实际上会不及格的考试。这项研究用具体数字量化了这种替代造成的能力虚报程度,对所有基于人类考试来评估AI的基准测试方法都有警示意义。
本文术语
- 凸性/阶梯式评分(convex marking scheme) · 得分不与答对数量成比例,而是在特定区间内跳跃或被压低的评分方式
- 部分得分差距(partial-credit gap/shortfall) · 按比例正确率算出的分数与实际阶梯评分之间的差值
- 百分位(percentile) · 表示某个分数超过了百分之多少真实考生的排名指标
- 开放权重模型(open-weight model) · 公开发布模型权重文件、任何人都可下载运行的AI模型
- Decision 764 · 越南教育培训部2025年发布的官方决定,规定了高中毕业考试的新题型和评分规则
论文原文摘要(英文)
When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an examination uses a non-additive grading scheme. The 2025 reform of Vietnam's National High School Graduation Examination demonstrates the cost of this substitution. In Part II of the exam, candidates eval
在 arXiv 阅读最新论文
- FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI SystemsAI经常能说对财务对账出错的原因,却拿不出真正的证据
- Looped Language Models Improve Compositional Tool Calling会反复回想自己答案的AI模型,更擅长按顺序组合调用多个工具
- FACET: Preserving Source Intent and Executable State in Terminal Task SynthesisFACET:让终端命令行任务的“说明书、环境、答案、判分器”自动保持一致
- FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents让AI连续管理一家足球俱乐部20年后发现,胜负关键不在模型大小,而在经营习惯
- Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-EngagementAI智能体追着离场用户发WhatsApp,把逛而不买的顾客拉回来
- Adversarial Review: Structured Disagreement for Grounded Agentic Code ReviewAI代码审查:与其堆更多智能体,不如让一个审查者和一个批评者互相较真
- GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence Networks把病毒基因序列变成密码子关系网络图,用来区分新冠变异株
- Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models语言模型全程冻结,只训练一个小连接器,也能做出好用的听觉理解AI