FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
让AI连续管理一家足球俱乐部20年后发现,胜负关键不在模型大小,而在经营习惯
FM-Bench是一个让语言模型智能体使用26种工具、在大约340到400个决策节点中管理一家虚拟足球俱乐部长达20个游戏年的基准测试。它设置了永远无法完全掌握的球员真实能力、要多年后才能见效的投资、会根据智能体行为调整报价的转会市场,以及同时考核战绩和财务纪律的董事会,以此检验模型能否在长期跨度中持续做出有效决策。在三次重复实验中,15个前沿模型分别在单人模式和多模型共享世界的竞技场模式下运行,claude-fable-5在两个赛道都夺得第一,但整体排名与模型规模、价格或厂商都没有关系。
他们做了什么
- 构建了一个长达20年的虚拟足球俱乐部经营环境,用来测试语言模型智能体在决策后果长期累积的情况下能否持续做出有效管理决策。
- 把现实经营中的四大难点做成具体机制:球员真实能力永远无法完全知晓、投资要多年后才见成效、转会市场会因被拒的报价而抬高隐藏价格、董事会同时考核战绩和财务纪律。
- 将15个前沿模型(来自Anthropic、OpenAI、Google、xAI、Meta的闭源模型,以及5个开源权重模型)分别投入单人赛道(对阵固定脚本世界)和竞技场赛道(15个模型共享同一个20年世界直接竞争)。
- 把最终总分拆解成六项行为能力指标,比如是否在临近结束时减少长周期投资、是否让现金闲置、是否提前开启合同续约谈判,以此分析哪些行为习惯真正决定成绩好坏。
- claude-fable-5在两个赛道都夺冠(单人赛道平均90.94分,而能看到隐藏信息的脚本参照策略为95.54分),但在竞技场中冠军头衔仍在十个不同模型之间轮换。
| demand | mechanism | targeted ability | calibration lever |
|---|---|---|---|
| Hidden information | scout bands with a permanent per-scout bias; hidden player traits (injury proneness, development rate, aging onset); hidden asks in negotiation | valuation calibration under noise | bandwidth, bias sigma |
| Cumulative consequences | Upside: youth development and facility investment pay off over years. Downside: insolvency ends in administration (points penalty, forced sales), confidence collapse ends in firing; both compound yet remain recoverable. Honors accrue into the final score season by season | long-range credit assignment; trend recognition, loss-cutting | growth and aging curves; spiral thresholds, warning cadence; honors weights |
| Counter-adaptive market | rejected bids raise the hidden ask; repeat-pair markups (bargains included); anti-inversion counters; negotiation cooldowns; per-seed mispricing | strategy adaptation | markup and decay rates |
| Multi-objective pressure | the board judges results and financial discipline jointly; season targets scale with squad strength | multi-constraint balancing | target cushion, board patience |
| seat | Sfinal | tokens |
|---|---|---|
| oracle (privileged) | 95.54±4.68 | — |
| claude-fable-5 | 90.94±5.20 | 24M |
| kimi-k2.6 | 88.49±0.15 | 87M |
| gpt-5.6-terra | 86.66±1.20 | 28M |
| gpt-5.6-sol | 86.40±2.53 | 62M |
| muse-spark-1.1 | 83.19±12.03 | 73M |
| glm-5.2 | 83.17±0.92 | 58M |
| grok-4.5 | 81.82±9.87 | 191M |
| qwen3.7-max | 80.66±10.38 | 47M |
| deepseek-v4-pro | 79.15±2.67 | 194M |
| gemini-3-flash | 79.06±13.45 | 51M |
| claude-sonnet-5 | 75.75±5.03 | 39M |
| claude-opus-4.8 | 75.02±2.46 | 31M |
| gemini-3.5-flash | 74.59±11.02 | 154M |
| minimax-m3 | 68.37±12.33 | 74M |
| claude-haiku-4.5 | 36.90±22.73 | 86M |
| heuristic | 17.05±12.34 | — |
| idle | −0.90±1.86 | — |
| random | −17.21±2.45 | — |
| player | Sfinal | settle | deaths | stops | time |
|---|---|---|---|---|---|
| H1 | 74.64 | completed | 0 | 394 | 10 h |
| H2 | 59.95 | completed | 0 | 398 | 4 h |
| H3 | 5.39 | fired, t=10.8 | 4 | 202 | 4 h |
| H4 | 2.47 | fired, t=10.7 | 4 | 198 | 5 h |
| H5 | 0.41 | fired, t=7.8 | 4 | 165 | 1.5 h |
| H6 | −3.78 | fired, t=12.3 | 4 | 237 | 2 h |
| # | seat | Sfinal | deaths | settle | tokens |
|---|---|---|---|---|---|
| 1 | claude-fable-5 | 76.26 | 0 | completed | ∼34M |
| 2 | muse-spark-1.1 | 62.47 | 0 | completed | ∼80M |
| 3 | deepseek-v4-pro | 52.09 | 0 | completed | ∼111M |
| 4 | glm-5.2 | 51.44 | 0 | completed | ∼86M |
| 5 | grok-4.5 | 50.89 | 0 | completed | ∼138M |
| 6 | gpt-5.6-sol | 47.94 | 0 | completed | ∼78M |
| 7 | minimax-m3 | 44.78 | 0 | completed | ∼65M |
| 8 | claude-sonnet-5 | 40.75 | 0 | completed | ∼42M |
| 9 | kimi-k2.6 | 39.72 | 0 | completed | ∼119M |
| 10 | gpt-5.6-terra | 38.44 | 0 | completed | ∼19M |
| 11 | qwen3.7-max | 32.77 | 2 | completed | ∼54M |
| 12 | claude-opus-4.8 | 28.44 | 1 | completed | ∼43M |
| 13 | gemini-3-flash | 15.76 | 2 | completed | ∼63M |
| 14 | claude-haiku-4.5 | 0.76 | 4 | fired, t=10.7 | ∼30M |
| 15 | gemini-3.5-flash | 0.21 | 4 | fired, t=5.8 | ∼40M |
| 16 | heuristic (anchor) | 0.13 | 4 | fired, t=2.6 | — |
| # | seat | provider | role |
|---|---|---|---|
| 1 | claude-fable-5 (4) | Anthropic | new flagship |
| 2 | claude-opus-4.8 (5) | Anthropic | previous flagship |
| 3 | claude-haiku-4.5 (3) | Anthropic | small tier |
| 4 | gpt-5.6-sol (32) | OpenAI | new flagship |
| 5 | gpt-5.6-terra (32) | OpenAI | balanced tier |
| 6 | claude-sonnet-5 (6) | Anthropic | workhorse mid (4-point family curve) |
| 7 | gemini-3.5-flash (16) | latest GA flagship | |
| 8 | gemini-3-flash (15) | previous-gen fast | |
| 9 | deepseek-v4-pro (40) | Together | open flagship |
| 10 | grok-4.5 (38) | xAI | frontier |
| 11 | kimi-k2.6 (30) | Together | open agentic specialist |
| 12 | qwen3.7-max (1) | Together | open flagship |
| 13 | glm-5.2 (42) | Together | open top tier |
| 14 | muse-spark-1.1 (27) | Meta Model API | Meta agentic model |
| 15 | minimax-m3 (29) | Together | open flagship |
| 16 | heuristic | — (scripted) | disciplined-script anchor |
| lever | easy | medium | hard |
|---|---|---|---|
| scouting noise, multiplier on the belief sigma | 1.3 | 1.0 | 0.35 |
| lineup rotates for fatigue | no | yes | yes |
| market days acted on | every 2nd | every | every |
| gap required before buying an upgrade | 1.5 | 1.0 | 0.25 |
| wage budget, multiplier on the board target | 0.92 | 1.0 | 1.10 |
| seat | score | sd | profile |
|---|---|---|---|
| claude-fable-5 | 90.9 | 5.2 | generalist, no axis below 0.79; best bidder in the field at 9 offers per signing and the sharpest endgame reduction |
| kimi-k2.6 | 88.5 | 0.1 | the steadiest model in the campaign, scoring within a third of a point on three different worlds |
| gpt-5.6-terra | 86.7 | 1.2 | minimalist: lowest token spend in the field (28M) after the winner and an early endgame reduction, but the shortest renewal leads among the top models |
| gpt-5.6-sol | 86.4 | 2.5 | archivist memory (similarity 0.91, append-only) and the weakest bidding of the top group (45 offers per signing) |
| muse-spark-1.1 | 83.2 | 12.0 | strong on two worlds and 20 points weaker on the third; one of two models that raise long-horizon spending at the end |
| glm-5.2 | 83.2 | 0.9 | stable across seeds but loose with cash (135% idle ratio) |
| grok-4.5 | 81.8 | 9.9 | compute as substitute: long renewal leads and heavy spend (191M tokens) with wide seed-to-seed swings |
| qwen3.7-max | 80.7 | 10.4 | longest renewal leads after the winner, undone by a 103% idle ratio and churning memory (similarity 0.23) |
| deepseek-v4-pro | 79.2 | 2.7 | near-static notebook (0.81) and the heaviest spend in the field (194M tokens) |
| gemini-3-flash | 79.1 | 13.5 | best cash discipline in the field (46%) but the least willing to cut endgame spending; collapses on the hardest seed |
| claude-sonnet-5 | 75.7 | 5.0 | churning memory (0.20, the lowest) and no endgame reduction |
| claude-opus-4.8 | 75.0 | 2.5 | shortest renewal leads in the field (10 months) with a 134% idle ratio |
| gemini-3.5-flash | 74.6 | 11.0 | worst price discovery (73 offers per signing) and the largest endgame ramp-up |
| minimax-m3 | 68.4 | 12.3 | prefers veteran signings, short leads, and swings 23 points across seeds |
| claude-haiku-4.5 | 36.9 | 22.7 | weak on every axis at once: 196% idle cash, 11-month leads, a rising endgame rate, and the widest spread in the campaign |
| seat | Y5 | Y10 | Y15 | Y20 |
|---|---|---|---|---|
| claude-fable-5 | 4th ⋅ 140H ⋅ 293M ⋅ 130M ⋅ 33.3 | 1st ⋅ 480H ⋅ 470M ⋅ 218M ⋅ 50.3 | 1st ⋅ 900H ⋅ 1048M ⋅ 443M ⋅ 65.9 | 1st ⋅ 1400H ⋅ 1742M ⋅ 634M ⋅ 94.6 |
| muse-spark-1.1 | 3rd ⋅ 120H ⋅ 302M ⋅ 158M ⋅ 31.1 | 1st ⋅ 460H ⋅ 468M ⋅ 196M ⋅ 48.7 | 1st ⋅ 960H ⋅ 1037M ⋅ 457M ⋅ 66.5 | 1st ⋅ 1460H ⋅ 1298M ⋅ 315M ⋅ 91.2 |
| grok-4.5 | 4th ⋅ 60H ⋅ 289M ⋅ 169M ⋅ 23.9 | 1st ⋅ 460H ⋅ 318M ⋅ 107M ⋅ 46.6 | 1st ⋅ 880H ⋅ 618M ⋅ 225M ⋅ 58.1 | 1st ⋅ 1380H ⋅ 1106M ⋅ 634M ⋅ 89.7 |
| kimi-k2.6 | 3rd ⋅ 120H ⋅ 344M ⋅ 116M ⋅ 31.9 | 1st ⋅ 440H ⋅ 430M ⋅ 99M ⋅ 46.7 | 1st ⋅ 940H ⋅ 668M ⋅ 276M ⋅ 61.4 | 1st ⋅ 1440H ⋅ 733M ⋅ 400M ⋅ 88.5 |
| gemini-3-flash | 3rd ⋅ 160H ⋅ 247M ⋅ 165M ⋅ 32.8 | 1st ⋅ 500H ⋅ 339M ⋅ 111M ⋅ 42.1 | 1st ⋅ 1000H ⋅ 675M ⋅ 400M ⋅ 61.5 | 1st ⋅ 1420H ⋅ 812M ⋅ 618M ⋅ 86.6 |
| qwen3.7-max | 3rd ⋅ 120H ⋅ 404M ⋅ 127M ⋅ 34.0 | 1st ⋅ 360H ⋅ 582M ⋅ 140M ⋅ 48.0 | 1st ⋅ 780H ⋅ 917M ⋅ 256M ⋅ 62.3 | 1st ⋅ 1200H ⋅ 1305M ⋅ 248M ⋅ 86.1 |
| gpt-5.6-terra | 3rd ⋅ 140H ⋅ 359M ⋅ 120M ⋅ 34.0 | 1st ⋅ 440H ⋅ 565M ⋅ 102M ⋅ 50.1 | 1st ⋅ 780H ⋅ 808M ⋅ 247M ⋅ 59.8 | 1st ⋅ 1280H ⋅ 987M ⋅ 186M ⋅ 85.3 |
| gpt-5.6-sol | 4th ⋅ 140H ⋅ 343M ⋅ 135M ⋅ 33.6 | 1st ⋅ 460H ⋅ 615M ⋅ 176M ⋅ 51.5 | 1st ⋅ 960H ⋅ 762M ⋅ 217M ⋅ 62.6 | 2nd ⋅ 1220H ⋅ 1109M ⋅ 274M ⋅ 84.5 |
| glm-5.2 | 3rd ⋅ 120H ⋅ 294M ⋅ 121M ⋅ 31.8 | 1st ⋅ 420H ⋅ 513M ⋅ 110M ⋅ 48.1 | 3rd ⋅ 760H ⋅ 712M ⋅ 58M ⋅ 56.6 | 1st ⋅ 1080H ⋅ 1331M ⋅ 266M ⋅ 82.8 |
| minimax-m3 | 10th ⋅ 20H ⋅ 324M ⋅ 154M ⋅ 20.9 | 1st ⋅ 240H ⋅ 492M ⋅ 100M ⋅ 39.3 | 1st ⋅ 540H ⋅ 827M ⋅ 155M ⋅ 51.8 | 1st ⋅ 1040H ⋅ 1245M ⋅ 327M ⋅ 82.5 |
| claude-sonnet-5 | 5th ⋅ 20H ⋅ 247M ⋅ 91M ⋅ 18.6 | 2nd ⋅ 180H ⋅ 363M ⋅ 97M ⋅ 36.0 | 1st ⋅ 480H ⋅ 647M ⋅ 94M ⋅ 50.2 | 1st ⋅ 980H ⋅ 1020M ⋅ 146M ⋅ 80.4 |
| deepseek-v4-pro | 5th ⋅ 140H ⋅ 385M ⋅ 124M ⋅ 35.8 | 1st ⋅ 480H ⋅ 638M ⋅ 111M ⋅ 52.9 | 1st ⋅ 820H ⋅ 732M ⋅ 110M ⋅ 58.7 | 3rd ⋅ 1080H ⋅ 996M ⋅ 163M ⋅ 79.0 |
| claude-opus-4.8 | 3rd ⋅ 140H ⋅ 308M ⋅ 100M ⋅ 32.7 | 1st ⋅ 280H ⋅ 467M ⋅ 78M ⋅ 40.5 | 2nd ⋅ 600H ⋅ 788M ⋅ 160M ⋅ 56.4 | 3rd ⋅ 760H ⋅ 1313M ⋅ 216M ⋅ 76.9 |
| gemini-3.5-flash | 2nd ⋅ 140H ⋅ 260M ⋅ 186M ⋅ 29.9 | 1st ⋅ 480H ⋅ 226M ⋅ 140M ⋅ 41.3 | 1st ⋅ 980H ⋅ 479M ⋅ 156M ⋅ 57.2 | 5th ⋅ 1040H ⋅ 877M ⋅ 170M ⋅ 75.2 |
| claude-haiku-4.5 | 12th ⋅ 100H ⋅ 293M ⋅ 132M ⋅ 28.4 | 7th ⋅ 200H ⋅ 395M ⋅ 39M ⋅ 33.8 | 3rd ⋅ 340H ⋅ 629M ⋅ 42M ⋅ 42.8 | 12th ⋅ 360H ⋅ 821M ⋅ 59M ⋅ 57.3 |
| seat | Y5 | Y10 | Y15 | Y20 |
|---|---|---|---|---|
| claude-fable-5 | 2nd ⋅ 180H ⋅ 326M ⋅ 105M ⋅ 47.4 | 5th ⋅ 240H ⋅ 394M ⋅ 196M ⋅ 56.4 | 5th ⋅ 440H ⋅ 568M ⋅ 250M ⋅ 71.3 | 3rd ⋅ 620H ⋅ 507M ⋅ 298M ⋅ 76.3 |
| muse-spark-1.1 | 8th ⋅ 0H ⋅ 260M ⋅ 84M ⋅ 18.5 | 7th ⋅ 40H ⋅ 197M ⋅ 117M ⋅ 24.0 | 2nd ⋅ 160H ⋅ 272M ⋅ 222M ⋅ 46.2 | 1st ⋅ 420H ⋅ 306M ⋅ 246M ⋅ 62.5 |
| deepseek-v4-pro | 9th ⋅ 0H ⋅ 279M ⋅ 91M ⋅ 20.0 | 10th ⋅ 140H ⋅ 342M ⋅ 91M ⋅ 44.5 | 4th ⋅ 200H ⋅ 397M ⋅ 122M ⋅ 52.2 | 5th ⋅ 200H ⋅ 363M ⋅ 160M ⋅ 52.1 |
| glm-5.2 | 4th ⋅ 40H ⋅ 275M ⋅ 107M ⋅ 30.3 | 1st ⋅ 160H ⋅ 335M ⋅ 78M ⋅ 46.1 | 8th ⋅ 200H ⋅ 346M ⋅ 71M ⋅ 49.3 | 6th ⋅ 200H ⋅ 356M ⋅ 125M ⋅ 51.4 |
| grok-4.5 | 7th ⋅ 20H ⋅ 197M ⋅ 134M ⋅ 21.1 | 6th ⋅ 260H ⋅ 145M ⋅ 145M ⋅ 40.0 | 7th ⋅ 320H ⋅ 108M ⋅ 185M ⋅ 36.4 | 8th ⋅ 520H ⋅ 139M ⋅ 199M ⋅ 50.9 |
| gpt-5.6-sol | 6th ⋅ 100H ⋅ 258M ⋅ 115M ⋅ 36.9 | 4th ⋅ 120H ⋅ 239M ⋅ 86M ⋅ 36.8 | 1st ⋅ 240H ⋅ 281M ⋅ 229M ⋅ 52.3 | 2nd ⋅ 280H ⋅ 182M ⋅ 294M ⋅ 47.9 |
| minimax-m3 | 3rd ⋅ 120H ⋅ 272M ⋅ 93M ⋅ 38.2 | 11th ⋅ 220H ⋅ 246M ⋅ 61M ⋅ 43.2 | 13th ⋅ 220H ⋅ 236M ⋅ 64M ⋅ 42.5 | 13th ⋅ 220H ⋅ 260M ⋅ 76M ⋅ 44.8 |
| claude-sonnet-5 | 11th ⋅ 100H ⋅ 246M ⋅ 82M ⋅ 35.9 | 13th ⋅ 100H ⋅ 268M ⋅ 64M ⋅ 36.4 | 9th ⋅ 100H ⋅ 251M ⋅ 77M ⋅ 36.0 | 10th ⋅ 160H ⋅ 226M ⋅ 94M ⋅ 40.8 |
| kimi-k2.6 | 13th ⋅ 0H ⋅ 238M ⋅ 76M ⋅ 18.1 | 8th ⋅ 20H ⋅ 305M ⋅ 82M ⋅ 26.9 | 3rd ⋅ 60H ⋅ 351M ⋅ 106M ⋅ 36.9 | 7th ⋅ 80H ⋅ 304M ⋅ 175M ⋅ 39.7 |
| gpt-5.6-terra | 5th ⋅ 40H ⋅ 266M ⋅ 108M ⋅ 28.8 | 2nd ⋅ 60H ⋅ 113M ⋅ 108M ⋅ 13.0 | 6th ⋅ 60H ⋅ 352M ⋅ 108M ⋅ 36.2 | 9th ⋅ 60H ⋅ 390M ⋅ 138M ⋅ 38.4 |
| qwen3.7-max | 14th ⋅ 20H ⋅ 279M ⋅ 87M ⋅ 19.2 | 12th ⋅ 20H ⋅ 254M ⋅ 88M ⋅ 13.6 | 12th ⋅ 180H ⋅ 451M ⋅ 200M ⋅ 34.2 | 11th ⋅ 180H ⋅ 451M ⋅ 113M ⋅ 32.8 |
| claude-opus-4.8 | 10th ⋅ 20H ⋅ 274M ⋅ 90M ⋅ 18.4 | 3rd ⋅ 40H ⋅ 263M ⋅ 99M ⋅ 21.2 | 10th ⋅ 40H ⋅ 412M ⋅ 132M ⋅ 27.8 | 14th ⋅ 40H ⋅ 472M ⋅ 95M ⋅ 28.4 |
| gemini-3-flash | 1st ⋅ 140H ⋅ −6M ⋅ 110M ⋅ 12.7 | 9th ⋅ 160H ⋅ −77M ⋅ 107M ⋅ 10.7 | 11th ⋅ 160H ⋅ 92M ⋅ 111M ⋅ 11.7 | 4th ⋅ 200H ⋅ 107M ⋅ 182M ⋅ 15.8 |
| claude-haiku-4.5 | 16th ⋅ 0H ⋅ 209M ⋅ 64M ⋅ 11.0 | 14th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 4.3 | 14th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 0.8 | 16th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 0.8 |
| gemini-3.5-flash | 12th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 2.1 | 15th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.2 | 15th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.2 | 12th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.2 |
| tool | what it does |
|---|---|
| Query (12), charged to the 30-query budget: | |
| get_club_overview | the agent’s club at a glance: division, cash, board, facilities, tactics |
| get_squad | the senior squad with ability bands, form, contracts, wages |
| get_player | detailed view of any player by id |
| get_league_table | league standings |
| get_fixtures | the club’s season fixtures and results |
| get_match_report | match report by match id |
| get_transfer_market | players available to buy: listed, free agents, expiring contracts |
| get_finances | cash, revenue, wage bill, active warnings |
| get_youth_academy | academy prospects with potential stars |
| get_history | the run’s archive: honors, seasons, transfers |
| get_inbox | pending offers and open items |
| get_draft_pool | draft phase only: the shared template pool with scouted ability bands, potential stars, prices, and preset contracts, plus budget status |
| Action (11); negotiation moves charged to the 10-move budget: | |
| set_lineup | set the preferred starting XI (auto-completed if players are unavailable) |
| set_tactics | set formation and playing style from fixed enums |
| make_transfer_offer | bid for another club’s player; the negotiation resolves within the stop |
| respond_to_offer | accept, reject, counter, or accept a counter on a pending offer |
| offer_contract | renew an own player or sign a free agent at a wage and term |
| list_player | put a player on or off the transfer list |
| release_player | terminate a contract (severance: half of one year’s wage) |
| promote_youth | promote an academy prospect to the senior squad |
| invest | start a facility upgrade: academy, training, or stadium |
| set_standing_order | add or clear a standing order (max 10; one threshold plus a fixed action) |
| submit_draft | draft phase only: submit the complete pick list; resubmitting replaces it, and auto-fill completes a short list from the cheapest tier |
| Notebook (2), free: | |
| append_note | append to the private notebook (persists across stops) |
| rewrite_notes | replace the entire notebook |
| Control (1): | |
| advance | end this decision stop and simulate to the next one |
| stop (day of the 360-day year) | what the world wants decided |
|---|---|
| Scheduled, 13 per season: | |
| Preseason board (2) | the board sets the season target; plan the year |
| Summer window plan (5) | the transfer window opens; build the squad |
| Summer deadline (59) | last actions before the window closes |
| Lineup lock (63) | commit lineup and tactics before round 1 |
| Monthly report (90, 120, 150, 240) | finances and form check-in |
| Winter window open (181) | mid-course squad correction |
| Midseason review (190) | the board measures progress against target |
| Winter deadline (209) | last winter-window actions |
| Season settlement (300) | the season closes: honors, revenue, board verdict |
| Youth intake (305) | the academy class arrives; promote, hold, or release |
| Event-driven, unsolicited: | |
| Draft (once, at run start) | assemble the 25-player squad from the shared pool |
| Offer batch | incoming bids for the agent’s players, batched at a minimum 15-day spacing |
| Contract expiry | a senior contract is running down; renew or lose the player for free |
| Major injury | a first-team player goes down; cover or reshuffle |
| Board warning | confidence is deteriorating; the board expects a response |
| Insolvency warning | cash runway is shrinking (30- and 60-day warnings) |
| Administration | insolvency executed: points penalty and forced sales |
| Revival | the seat restarts under the capped-revival mechanism (§4.2) |
| layer | who writes it | contents | delivery |
|---|---|---|---|
| Packet auto-context | engine | dashboard digest, inbox (all events since last stop), agent’s last 10 engine-mutating actions, pending offers | pushed free into every stop packet |
| Archive | engine | honors / season / transfer history of the whole run | get_history query tool, on demand |
| Notebook | the agent | intentions, reasoning, plans (“why I bought him”, “sell Y in winter”) | injected into every packet; written via append_note / rewrite_notes |
| Audit log | runner | every successful action | never fed back: replay/anti-cheat/resume only |
| quantity | median | range |
|---|---|---|
| decision stops | 374 | 341–390 |
| of which scheduled | ∼260 (13/season) | rest event-triggered (Table 11) |
| matches simulated in the world | 4,800 | (30 rounds × 8 matches × 20 seasons) |
| model turns (API calls) | ∼1,760 | 1,184–6,969 |
| state-changing actions | — | 849–2,198 |
| tokens (manifest accounting) | 61M | 24M–274M |
| notebook writes | 354 | 320–425 |
| notebook size at year 20 | — | 0.4k–209k characters |
| simulation parameters (frozen) | 313 | — |
为什么重要
这项工作把智能体是否具备长期经营能力这件事变得可以量化测量,更贴近现实中组织运作依赖持续、累积决策的方式。对于想把智能体用于公司运营、长期投资、资源管理等实际长期任务的人来说,它给出的结论是应该关注行为习惯而非模型规模或价格标签。
本文术语
- 智能体(agent) · 能自主选择并调用工具来完成任务的AI系统
- 单人赛道/竞技场 · 单人赛道是模型独自对阵固定脚本世界,竞技场是多个模型在同一个共享世界中直接竞争
- oracle(全知策略) · 拥有查看隐藏信息特权的脚本参照策略,用作性能的软上限基准
- 记事本(notebook) · 智能体每次决策时对话都会重置,记事本是它唯一能自己留存并传递给未来自己的信息载体
- Sfinal(最终得分) · 综合荣誉积分、净资产增值和阵容价值,在20年全程内累积计算出的单一总分
无法转载的图表
- Figure 1: How a run works. One run spans 20 in-game years and roughly 340 to 400 decision stops, ending in a single composite score. At each stop, the clock freezes and the agent takes as many tool-call turns as it wants (read → act → advance). The agent’s only cross-stop memory is its self-written notebook (§3.2); the environment side holds the club (net worth, squad with permanently hidden true ability, a board that can fire the manager) and the world (15 rival clubs, a counter-adaptive market, delayed-payoff investments).
- Figure 2: Per-season solo trajectories for all 15 models in four channels. Lines are the mean over three seeds, shading spans the best and worst seed, and the five models with the highest mean are colored and named in the legend; the grey lines are the remaining ten models.
- Figure 3: Credit assignment. Left: discretionary spend bucketed by payoff horizon, with run totals at right. Right: mean season-end idle-cash ratio vs. final score, rs=−0.50, negative on every seed.
- Figure 4: Proactive control. Left: distribution of contract months remaining when each renewal negotiation episode opens (rows sorted by final score; shaded band = last-minute, ≤6 months). Right: per-model median lead vs. final score, rs=+0.45, positive on every seed.
- Figure 5: Capability matrix over the 15 solo models, six axes. Each column rank-normalizes one behavioral metric averaged over the three seeds (1 = best of 15; construction in Appendix 15), and rows are sorted by mean final score. The winner is a generalist, with no axis below 0.79 and mean 0.94, rather than the leader of any single one; the mid-table pairs a genuine strength with a decisive gap; and the bottom of the board is weak on most axes at once.
- Figure 6: Arena trajectories, all 16 seats. The six highest-scoring seats are colored and named in the legend; the grey lines are the remaining ten seats. Left: composite Sfinal at each season end. Right: league-table position.
- Figure 7: Final score per model, mean with standard deviation over the three seeds, individual seeds shown as points. Dashed and dotted lines mark the oracle and heuristic means. Stability separates models that a single-seed board cannot tell apart, from kimi-k2.6 at 0.15 to claude-haiku-4.5 at 22.73.
- Figure 8: Transfer offers per season, solo track (rows sorted by final score; totals at right). The winner’s 9 offers are all in the first two seasons; the field-wide offer-to-completion conversion is 2–5%.
- Figure 9: Transfer offers per season in the Arena (approximate year mapping; daggered models settled early). The winner’s offer count rises from 9 (solo) to 91: the same model under a different, correct judgment of market liquidity.
- Figure 10: Season-by-season cumulative API cost vs. composite score, solo track. Steep curves convert dollars into score throughout; flat long curves do not. The full 20-year runs span $18 to $191.
- Figure 11: The failed metric: share of stops carrying an engine warning, by type (left) and against final score (right), means over the three seeds. The per-seed coefficient runs +0.45 / +0.07 / −0.78 and averages to nothing, so warning exposure is not reported as a capability. Zero warnings conflates genuine anticipation with do-nothing conservatism.
- Figure 12: Both panels are means over the three seeds. Left: youth harvest rate (share of promotions later fielded in ≥5 lineups); the correlation that looked informative on seed 1 does not survive three (−0.52 / −0.18 / +0.04). Right: endgame shift in long-horizon actions per season, years 2–16 (blue) vs. 17–20 (red); endgame-aware models reduce late investment, one ramps up.
- Figure 13: Memory-curation regimes on the solo track. Left: season-to-season TF–IDF cosine similarity of each model’s reconstructed season-end notebook. Right: notebook size over the run. The right panel decodes the left: the same “high consistency” is a 200k-character append-only archive for one model and a 3–6k curated document for the winner.
- Figure 14: Compute allocation. Mean tokens per run against mean solo score, both over the three seeds, on a log axis. A null result (rs=−0.19, p=0.50) across a sevenfold spend range.
- Figure 15: The asset channels behind Figure 6: per-season net worth and squad value for all 16 Arena seats, the six highest-scoring seats colored as in Figure 6 and the grey lines the remaining ten (numeric snapshots in Table 9). The winner’s mid-table seasons are asset accumulation (the net-worth curve keeps climbing while the standings do not), and the conservative cash-holders’ net worth grows on cash the composite discounts.
论文原文摘要(英文)
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget a
在 arXiv 阅读最新论文
- FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI SystemsAI经常能说对财务对账出错的原因,却拿不出真正的证据
- Looped Language Models Improve Compositional Tool Calling会反复回想自己答案的AI模型,更擅长按顺序组合调用多个工具
- FACET: Preserving Source Intent and Executable State in Terminal Task SynthesisFACET:让终端命令行任务的“说明书、环境、答案、判分器”自动保持一致
- Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-EngagementAI智能体追着离场用户发WhatsApp,把逛而不买的顾客拉回来
- Adversarial Review: Structured Disagreement for Grounded Agentic Code ReviewAI代码审查:与其堆更多智能体,不如让一个审查者和一个批评者互相较真
- GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence Networks把病毒基因序列变成密码子关系网络图,用来区分新冠变异株
- Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models语言模型全程冻结,只训练一个小连接器,也能做出好用的听觉理解AI
- Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval图文检索AI误以为自己已经全学会了,结果学不会区分那些几乎一样的描述句子