每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

arXiv:2608.184232026-08-20

让AI连续管理一家足球俱乐部20年后发现,胜负关键不在模型大小,而在经营习惯

FM-Bench是一个让语言模型智能体使用26种工具、在大约340到400个决策节点中管理一家虚拟足球俱乐部长达20个游戏年的基准测试。它设置了永远无法完全掌握的球员真实能力、要多年后才能见效的投资、会根据智能体行为调整报价的转会市场,以及同时考核战绩和财务纪律的董事会,以此检验模型能否在长期跨度中持续做出有效决策。在三次重复实验中,15个前沿模型分别在单人模式和多模型共享世界的竞技场模式下运行,claude-fable-5在两个赛道都夺得第一,但整体排名与模型规模、价格或厂商都没有关系。

他们做了什么

  1. 构建了一个长达20年的虚拟足球俱乐部经营环境,用来测试语言模型智能体在决策后果长期累积的情况下能否持续做出有效管理决策。
  2. 把现实经营中的四大难点做成具体机制:球员真实能力永远无法完全知晓、投资要多年后才见成效、转会市场会因被拒的报价而抬高隐藏价格、董事会同时考核战绩和财务纪律。
  3. 将15个前沿模型(来自Anthropic、OpenAI、Google、xAI、Meta的闭源模型,以及5个开源权重模型)分别投入单人赛道(对阵固定脚本世界)和竞技场赛道(15个模型共享同一个20年世界直接竞争)。
  4. 把最终总分拆解成六项行为能力指标,比如是否在临近结束时减少长周期投资、是否让现金闲置、是否提前开启合同续约谈判,以此分析哪些行为习惯真正决定成绩好坏。
  5. claude-fable-5在两个赛道都夺冠(单人赛道平均90.94分,而能看到隐藏信息的脚本参照策略为95.54分),但在竞技场中冠军头衔仍在十个不同模型之间轮换。
Table 1: The four demands: the mechanism that instantiates each, the ability it targets, and the calibration lever that tunes it.
demandmechanismtargeted abilitycalibration lever
Hidden informationscout bands with a permanent per-scout bias; hidden player traits (injury proneness, development rate, aging onset); hidden asks in negotiationvaluation calibration under noisebandwidth, bias sigma
Cumulative consequencesUpside: youth development and facility investment pay off over years. Downside: insolvency ends in administration (points penalty, forced sales), confidence collapse ends in firing; both compound yet remain recoverable. Honors accrue into the final score season by seasonlong-range credit assignment; trend recognition, loss-cuttinggrowth and aging curves; spiral thresholds, warning cadence; honors weights
Counter-adaptive marketrejected bids raise the hidden ask; repeat-pair markups (bargains included); anti-inversion counters; negotiation cooldowns; per-seed mispricingstrategy adaptationmarkup and decay rates
Multi-objective pressurethe board judges results and financial discipline jointly; season targets scale with squad strengthmulti-constraint balancingtarget cushion, board patience
Table 2: The Solo track results. The four non-LLM seats are reference policies. The oracle reads true hidden state through the same interface (§4.1, Appendix 13); heuristic is a disciplined blind script, idle never acts, and random acts arbitrarily (Appendix 10).
seatSfinaltokens
oracle (privileged)95.54±4.68
claude-fable-590.94±5.2024M
kimi-k2.688.49±0.1587M
gpt-5.6-terra86.66±1.2028M
gpt-5.6-sol86.40±2.5362M
muse-spark-1.183.19±12.0373M
glm-5.283.17±0.9258M
grok-4.581.82±9.87191M
qwen3.7-max80.66±10.3847M
deepseek-v4-pro79.15±2.67194M
gemini-3-flash79.06±13.4551M
claude-sonnet-575.75±5.0339M
claude-opus-4.875.02±2.4631M
gemini-3.5-flash74.59±11.02154M
minimax-m368.37±12.3374M
claude-haiku-4.536.90±22.7386M
heuristic17.05±12.34
idle−0.90±1.86
random−17.21±2.45
Table 3: The six first-play human runs (§4.2.2). Deaths are capped revivals spent before the settling firing; t is the game-year at which a fired run settled.
playerSfinalsettledeathsstopstime
H174.64completed039410 h
H259.95completed03984 h
H35.39fired, t=10.842024 h
H42.47fired, t=10.741985 h
H50.41fired, t=7.841651.5 h
H6−3.78fired, t=12.342372 h
Table 4: The Arena results.
#seatSfinaldeathssettletokens
1claude-fable-576.260completed∼34M
2muse-spark-1.162.470completed∼80M
3deepseek-v4-pro52.090completed∼111M
4glm-5.251.440completed∼86M
5grok-4.550.890completed∼138M
6gpt-5.6-sol47.940completed∼78M
7minimax-m344.780completed∼65M
8claude-sonnet-540.750completed∼42M
9kimi-k2.639.720completed∼119M
10gpt-5.6-terra38.440completed∼19M
11qwen3.7-max32.772completed∼54M
12claude-opus-4.828.441completed∼43M
13gemini-3-flash15.762completed∼63M
14claude-haiku-4.50.764fired, t=10.7∼30M
15gemini-3.5-flash0.214fired, t=5.8∼40M
16heuristic (anchor)0.134fired, t=2.6
Table 5: The roster: 15 frontier LLM seats plus one scripted anchor, July 2026.
#seatproviderrole
1claude-fable-5 (4)Anthropicnew flagship
2claude-opus-4.8 (5)Anthropicprevious flagship
3claude-haiku-4.5 (3)Anthropicsmall tier
4gpt-5.6-sol (32)OpenAInew flagship
5gpt-5.6-terra (32)OpenAIbalanced tier
6claude-sonnet-5 (6)Anthropicworkhorse mid (4-point family curve)
7gemini-3.5-flash (16)Googlelatest GA flagship
8gemini-3-flash (15)Googleprevious-gen fast
9deepseek-v4-pro (40)Togetheropen flagship
10grok-4.5 (38)xAIfrontier
11kimi-k2.6 (30)Togetheropen agentic specialist
12qwen3.7-max (1)Togetheropen flagship
13glm-5.2 (42)Togetheropen top tier
14muse-spark-1.1 (27)Meta Model APIMeta agentic model
15minimax-m3 (29)Togetheropen flagship
16heuristic— (scripted)disciplined-script anchor
Table 6: The five competence levers. Medium is the neutral setting.
levereasymediumhard
scouting noise, multiplier on the belief sigma1.31.00.35
lineup rotates for fatiguenoyesyes
market days acted onevery 2ndeveryevery
gap required before buying an upgrade1.51.00.25
wage budget, multiplier on the board target0.921.01.10
Table 7: Per-model profiles over the three seeds. Score is the mean and SD the standard deviation of Sfinal; axis extremes come from the matrix of Figure 5.
seatscoresdprofile
claude-fable-590.95.2generalist, no axis below 0.79; best bidder in the field at 9 offers per signing and the sharpest endgame reduction
kimi-k2.688.50.1the steadiest model in the campaign, scoring within a third of a point on three different worlds
gpt-5.6-terra86.71.2minimalist: lowest token spend in the field (28M) after the winner and an early endgame reduction, but the shortest renewal leads among the top models
gpt-5.6-sol86.42.5archivist memory (similarity 0.91, append-only) and the weakest bidding of the top group (45 offers per signing)
muse-spark-1.183.212.0strong on two worlds and 20 points weaker on the third; one of two models that raise long-horizon spending at the end
glm-5.283.20.9stable across seeds but loose with cash (135% idle ratio)
grok-4.581.89.9compute as substitute: long renewal leads and heavy spend (191M tokens) with wide seed-to-seed swings
qwen3.7-max80.710.4longest renewal leads after the winner, undone by a 103% idle ratio and churning memory (similarity 0.23)
deepseek-v4-pro79.22.7near-static notebook (0.81) and the heaviest spend in the field (194M tokens)
gemini-3-flash79.113.5best cash discipline in the field (46%) but the least willing to cut endgame spending; collapses on the hardest seed
claude-sonnet-575.75.0churning memory (0.20, the lowest) and no endgame reduction
claude-opus-4.875.02.5shortest renewal leads in the field (10 months) with a 134% idle ratio
gemini-3.5-flash74.611.0worst price discovery (73 offers per signing) and the largest endgame ramp-up
minimax-m368.412.3prefers veteran signings, short leads, and swings 23 points across seeds
claude-haiku-4.536.922.7weak on every axis at once: 196% idle cash, 11-month leads, a rising endgame rate, and the widest spread in the campaign
Table 8: Solo season snapshots (seed 1). Cell format: position ⋅ honors ⋅ net worth ⋅ squad value ⋅ score. Rows sorted by final score.
seatY5Y10Y15Y20
claude-fable-54th ⋅ 140H ⋅ 293M ⋅ 130M ⋅ 33.31st ⋅ 480H ⋅ 470M ⋅ 218M ⋅ 50.31st ⋅ 900H ⋅ 1048M ⋅ 443M ⋅ 65.91st ⋅ 1400H ⋅ 1742M ⋅ 634M ⋅ 94.6
muse-spark-1.13rd ⋅ 120H ⋅ 302M ⋅ 158M ⋅ 31.11st ⋅ 460H ⋅ 468M ⋅ 196M ⋅ 48.71st ⋅ 960H ⋅ 1037M ⋅ 457M ⋅ 66.51st ⋅ 1460H ⋅ 1298M ⋅ 315M ⋅ 91.2
grok-4.54th ⋅ 60H ⋅ 289M ⋅ 169M ⋅ 23.91st ⋅ 460H ⋅ 318M ⋅ 107M ⋅ 46.61st ⋅ 880H ⋅ 618M ⋅ 225M ⋅ 58.11st ⋅ 1380H ⋅ 1106M ⋅ 634M ⋅ 89.7
kimi-k2.63rd ⋅ 120H ⋅ 344M ⋅ 116M ⋅ 31.91st ⋅ 440H ⋅ 430M ⋅ 99M ⋅ 46.71st ⋅ 940H ⋅ 668M ⋅ 276M ⋅ 61.41st ⋅ 1440H ⋅ 733M ⋅ 400M ⋅ 88.5
gemini-3-flash3rd ⋅ 160H ⋅ 247M ⋅ 165M ⋅ 32.81st ⋅ 500H ⋅ 339M ⋅ 111M ⋅ 42.11st ⋅ 1000H ⋅ 675M ⋅ 400M ⋅ 61.51st ⋅ 1420H ⋅ 812M ⋅ 618M ⋅ 86.6
qwen3.7-max3rd ⋅ 120H ⋅ 404M ⋅ 127M ⋅ 34.01st ⋅ 360H ⋅ 582M ⋅ 140M ⋅ 48.01st ⋅ 780H ⋅ 917M ⋅ 256M ⋅ 62.31st ⋅ 1200H ⋅ 1305M ⋅ 248M ⋅ 86.1
gpt-5.6-terra3rd ⋅ 140H ⋅ 359M ⋅ 120M ⋅ 34.01st ⋅ 440H ⋅ 565M ⋅ 102M ⋅ 50.11st ⋅ 780H ⋅ 808M ⋅ 247M ⋅ 59.81st ⋅ 1280H ⋅ 987M ⋅ 186M ⋅ 85.3
gpt-5.6-sol4th ⋅ 140H ⋅ 343M ⋅ 135M ⋅ 33.61st ⋅ 460H ⋅ 615M ⋅ 176M ⋅ 51.51st ⋅ 960H ⋅ 762M ⋅ 217M ⋅ 62.62nd ⋅ 1220H ⋅ 1109M ⋅ 274M ⋅ 84.5
glm-5.23rd ⋅ 120H ⋅ 294M ⋅ 121M ⋅ 31.81st ⋅ 420H ⋅ 513M ⋅ 110M ⋅ 48.13rd ⋅ 760H ⋅ 712M ⋅ 58M ⋅ 56.61st ⋅ 1080H ⋅ 1331M ⋅ 266M ⋅ 82.8
minimax-m310th ⋅ 20H ⋅ 324M ⋅ 154M ⋅ 20.91st ⋅ 240H ⋅ 492M ⋅ 100M ⋅ 39.31st ⋅ 540H ⋅ 827M ⋅ 155M ⋅ 51.81st ⋅ 1040H ⋅ 1245M ⋅ 327M ⋅ 82.5
claude-sonnet-55th ⋅ 20H ⋅ 247M ⋅ 91M ⋅ 18.62nd ⋅ 180H ⋅ 363M ⋅ 97M ⋅ 36.01st ⋅ 480H ⋅ 647M ⋅ 94M ⋅ 50.21st ⋅ 980H ⋅ 1020M ⋅ 146M ⋅ 80.4
deepseek-v4-pro5th ⋅ 140H ⋅ 385M ⋅ 124M ⋅ 35.81st ⋅ 480H ⋅ 638M ⋅ 111M ⋅ 52.91st ⋅ 820H ⋅ 732M ⋅ 110M ⋅ 58.73rd ⋅ 1080H ⋅ 996M ⋅ 163M ⋅ 79.0
claude-opus-4.83rd ⋅ 140H ⋅ 308M ⋅ 100M ⋅ 32.71st ⋅ 280H ⋅ 467M ⋅ 78M ⋅ 40.52nd ⋅ 600H ⋅ 788M ⋅ 160M ⋅ 56.43rd ⋅ 760H ⋅ 1313M ⋅ 216M ⋅ 76.9
gemini-3.5-flash2nd ⋅ 140H ⋅ 260M ⋅ 186M ⋅ 29.91st ⋅ 480H ⋅ 226M ⋅ 140M ⋅ 41.31st ⋅ 980H ⋅ 479M ⋅ 156M ⋅ 57.25th ⋅ 1040H ⋅ 877M ⋅ 170M ⋅ 75.2
claude-haiku-4.512th ⋅ 100H ⋅ 293M ⋅ 132M ⋅ 28.47th ⋅ 200H ⋅ 395M ⋅ 39M ⋅ 33.83rd ⋅ 340H ⋅ 629M ⋅ 42M ⋅ 42.812th ⋅ 360H ⋅ 821M ⋅ 59M ⋅ 57.3
Table 9: Arena season snapshots (seed 7, one shared world); the 15 models, as on the public page (the scripted anchor settled at t≈2.6y). Cell format as in Table 8. Seats that exhausted the revival cap keep their frozen score afterward, visible as repeated values (claude-haiku-4.5 from Y15, gemini-3.5-flash from Y10). Rows sorted by final score.
seatY5Y10Y15Y20
claude-fable-52nd ⋅ 180H ⋅ 326M ⋅ 105M ⋅ 47.45th ⋅ 240H ⋅ 394M ⋅ 196M ⋅ 56.45th ⋅ 440H ⋅ 568M ⋅ 250M ⋅ 71.33rd ⋅ 620H ⋅ 507M ⋅ 298M ⋅ 76.3
muse-spark-1.18th ⋅ 0H ⋅ 260M ⋅ 84M ⋅ 18.57th ⋅ 40H ⋅ 197M ⋅ 117M ⋅ 24.02nd ⋅ 160H ⋅ 272M ⋅ 222M ⋅ 46.21st ⋅ 420H ⋅ 306M ⋅ 246M ⋅ 62.5
deepseek-v4-pro9th ⋅ 0H ⋅ 279M ⋅ 91M ⋅ 20.010th ⋅ 140H ⋅ 342M ⋅ 91M ⋅ 44.54th ⋅ 200H ⋅ 397M ⋅ 122M ⋅ 52.25th ⋅ 200H ⋅ 363M ⋅ 160M ⋅ 52.1
glm-5.24th ⋅ 40H ⋅ 275M ⋅ 107M ⋅ 30.31st ⋅ 160H ⋅ 335M ⋅ 78M ⋅ 46.18th ⋅ 200H ⋅ 346M ⋅ 71M ⋅ 49.36th ⋅ 200H ⋅ 356M ⋅ 125M ⋅ 51.4
grok-4.57th ⋅ 20H ⋅ 197M ⋅ 134M ⋅ 21.16th ⋅ 260H ⋅ 145M ⋅ 145M ⋅ 40.07th ⋅ 320H ⋅ 108M ⋅ 185M ⋅ 36.48th ⋅ 520H ⋅ 139M ⋅ 199M ⋅ 50.9
gpt-5.6-sol6th ⋅ 100H ⋅ 258M ⋅ 115M ⋅ 36.94th ⋅ 120H ⋅ 239M ⋅ 86M ⋅ 36.81st ⋅ 240H ⋅ 281M ⋅ 229M ⋅ 52.32nd ⋅ 280H ⋅ 182M ⋅ 294M ⋅ 47.9
minimax-m33rd ⋅ 120H ⋅ 272M ⋅ 93M ⋅ 38.211th ⋅ 220H ⋅ 246M ⋅ 61M ⋅ 43.213th ⋅ 220H ⋅ 236M ⋅ 64M ⋅ 42.513th ⋅ 220H ⋅ 260M ⋅ 76M ⋅ 44.8
claude-sonnet-511th ⋅ 100H ⋅ 246M ⋅ 82M ⋅ 35.913th ⋅ 100H ⋅ 268M ⋅ 64M ⋅ 36.49th ⋅ 100H ⋅ 251M ⋅ 77M ⋅ 36.010th ⋅ 160H ⋅ 226M ⋅ 94M ⋅ 40.8
kimi-k2.613th ⋅ 0H ⋅ 238M ⋅ 76M ⋅ 18.18th ⋅ 20H ⋅ 305M ⋅ 82M ⋅ 26.93rd ⋅ 60H ⋅ 351M ⋅ 106M ⋅ 36.97th ⋅ 80H ⋅ 304M ⋅ 175M ⋅ 39.7
gpt-5.6-terra5th ⋅ 40H ⋅ 266M ⋅ 108M ⋅ 28.82nd ⋅ 60H ⋅ 113M ⋅ 108M ⋅ 13.06th ⋅ 60H ⋅ 352M ⋅ 108M ⋅ 36.29th ⋅ 60H ⋅ 390M ⋅ 138M ⋅ 38.4
qwen3.7-max14th ⋅ 20H ⋅ 279M ⋅ 87M ⋅ 19.212th ⋅ 20H ⋅ 254M ⋅ 88M ⋅ 13.612th ⋅ 180H ⋅ 451M ⋅ 200M ⋅ 34.211th ⋅ 180H ⋅ 451M ⋅ 113M ⋅ 32.8
claude-opus-4.810th ⋅ 20H ⋅ 274M ⋅ 90M ⋅ 18.43rd ⋅ 40H ⋅ 263M ⋅ 99M ⋅ 21.210th ⋅ 40H ⋅ 412M ⋅ 132M ⋅ 27.814th ⋅ 40H ⋅ 472M ⋅ 95M ⋅ 28.4
gemini-3-flash1st ⋅ 140H ⋅ −6M ⋅ 110M ⋅ 12.79th ⋅ 160H ⋅ −77M ⋅ 107M ⋅ 10.711th ⋅ 160H ⋅ 92M ⋅ 111M ⋅ 11.74th ⋅ 200H ⋅ 107M ⋅ 182M ⋅ 15.8
claude-haiku-4.516th ⋅ 0H ⋅ 209M ⋅ 64M ⋅ 11.014th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 4.314th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 0.816th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 0.8
gemini-3.5-flash12th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 2.115th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.215th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.212th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.2
Table 10: The 26-tool interface: 12 query tools, 11 action tools, 2 notebook tools, and 1 control tool. Descriptions follow the tool schemas served to the agent.
toolwhat it does
Query (12), charged to the 30-query budget:
get_club_overviewthe agent’s club at a glance: division, cash, board, facilities, tactics
get_squadthe senior squad with ability bands, form, contracts, wages
get_playerdetailed view of any player by id
get_league_tableleague standings
get_fixturesthe club’s season fixtures and results
get_match_reportmatch report by match id
get_transfer_marketplayers available to buy: listed, free agents, expiring contracts
get_financescash, revenue, wage bill, active warnings
get_youth_academyacademy prospects with potential stars
get_historythe run’s archive: honors, seasons, transfers
get_inboxpending offers and open items
get_draft_pooldraft phase only: the shared template pool with scouted ability bands, potential stars, prices, and preset contracts, plus budget status
Action (11); negotiation moves charged to the 10-move budget:
set_lineupset the preferred starting XI (auto-completed if players are unavailable)
set_tacticsset formation and playing style from fixed enums
make_transfer_offerbid for another club’s player; the negotiation resolves within the stop
respond_to_offeraccept, reject, counter, or accept a counter on a pending offer
offer_contractrenew an own player or sign a free agent at a wage and term
list_playerput a player on or off the transfer list
release_playerterminate a contract (severance: half of one year’s wage)
promote_youthpromote an academy prospect to the senior squad
investstart a facility upgrade: academy, training, or stadium
set_standing_orderadd or clear a standing order (max 10; one threshold plus a fixed action)
submit_draftdraft phase only: submit the complete pick list; resubmitting replaces it, and auto-fill completes a short list from the cheapest tier
Notebook (2), free:
append_noteappend to the private notebook (persists across stops)
rewrite_notesreplace the entire notebook
Control (1):
advanceend this decision stop and simulate to the next one
Table 11: The decision-stop calendar: 13 scheduled stops per season on the 360-day year, plus event-driven stops the world raises unsolicited. Several types can fire at one stop.
stop (day of the 360-day year)what the world wants decided
Scheduled, 13 per season:
Preseason board (2)the board sets the season target; plan the year
Summer window plan (5)the transfer window opens; build the squad
Summer deadline (59)last actions before the window closes
Lineup lock (63)commit lineup and tactics before round 1
Monthly report (90, 120, 150, 240)finances and form check-in
Winter window open (181)mid-course squad correction
Midseason review (190)the board measures progress against target
Winter deadline (209)last winter-window actions
Season settlement (300)the season closes: honors, revenue, board verdict
Youth intake (305)the academy class arrives; promote, hold, or release
Event-driven, unsolicited:
Draft (once, at run start)assemble the 25-player squad from the shared pool
Offer batchincoming bids for the agent’s players, batched at a minimum 15-day spacing
Contract expirya senior contract is running down; renew or lose the player for free
Major injurya first-team player goes down; cover or reshuffle
Board warningconfidence is deteriorating; the board expects a response
Insolvency warningcash runway is shrinking (30- and 60-day warnings)
Administrationinsolvency executed: points penalty and forced sales
Revivalthe seat restarts under the capped-revival mechanism (§4.2)
Table 12: The four memory layers (§3.2). What happened is recorded automatically; what it means and what to do next must be written by the agent itself.
layerwho writes itcontentsdelivery
Packet auto-contextenginedashboard digest, inbox (all events since last stop), agent’s last 10 engine-mutating actions, pending offerspushed free into every stop packet
Archiveenginehonors / season / transfer history of the whole runget_history query tool, on demand
Notebookthe agentintentions, reasoning, plans (“why I bought him”, “sell Y in winter”)injected into every packet; written via append_note / rewrite_notes
Audit logrunnerevery successful actionnever fed back: replay/anti-cheat/resume only
Table 13: One 20-year run, by the numbers: medians and ranges across the 15 solo models.
quantitymedianrange
decision stops374341–390
of which scheduled∼260 (13/season)rest event-triggered (Table 11)
matches simulated in the world4,800(30 rounds × 8 matches × 20 seasons)
model turns (API calls)∼1,7601,184–6,969
state-changing actions849–2,198
tokens (manifest accounting)61M24M–274M
notebook writes354320–425
notebook size at year 200.4k–209k characters
simulation parameters (frozen)313

为什么重要

这项工作把智能体是否具备长期经营能力这件事变得可以量化测量,更贴近现实中组织运作依赖持续、累积决策的方式。对于想把智能体用于公司运营、长期投资、资源管理等实际长期任务的人来说,它给出的结论是应该关注行为习惯而非模型规模或价格标签。

本文术语

  • 智能体(agent) · 能自主选择并调用工具来完成任务的AI系统
  • 单人赛道/竞技场 · 单人赛道是模型独自对阵固定脚本世界,竞技场是多个模型在同一个共享世界中直接竞争
  • oracle(全知策略) · 拥有查看隐藏信息特权的脚本参照策略,用作性能的软上限基准
  • 记事本(notebook) · 智能体每次决策时对话都会重置,记事本是它唯一能自己留存并传递给未来自己的信息载体
  • Sfinal(最终得分) · 综合荣誉积分、净资产增值和阵容价值,在20年全程内累积计算出的单一总分

无法转载的图表

  • Figure 1: How a run works. One run spans 20 in-game years and roughly 340 to 400 decision stops, ending in a single composite score. At each stop, the clock freezes and the agent takes as many tool-call turns as it wants (read → act → advance). The agent’s only cross-stop memory is its self-written notebook (§3.2); the environment side holds the club (net worth, squad with permanently hidden true ability, a board that can fire the manager) and the world (15 rival clubs, a counter-adaptive market, delayed-payoff investments).
  • Figure 2: Per-season solo trajectories for all 15 models in four channels. Lines are the mean over three seeds, shading spans the best and worst seed, and the five models with the highest mean are colored and named in the legend; the grey lines are the remaining ten models.
  • Figure 3: Credit assignment. Left: discretionary spend bucketed by payoff horizon, with run totals at right. Right: mean season-end idle-cash ratio vs. final score, rs=−0.50, negative on every seed.
  • Figure 4: Proactive control. Left: distribution of contract months remaining when each renewal negotiation episode opens (rows sorted by final score; shaded band = last-minute, ≤6 months). Right: per-model median lead vs. final score, rs=+0.45, positive on every seed.
  • Figure 5: Capability matrix over the 15 solo models, six axes. Each column rank-normalizes one behavioral metric averaged over the three seeds (1 = best of 15; construction in Appendix 15), and rows are sorted by mean final score. The winner is a generalist, with no axis below 0.79 and mean 0.94, rather than the leader of any single one; the mid-table pairs a genuine strength with a decisive gap; and the bottom of the board is weak on most axes at once.
  • Figure 6: Arena trajectories, all 16 seats. The six highest-scoring seats are colored and named in the legend; the grey lines are the remaining ten seats. Left: composite Sfinal at each season end. Right: league-table position.
  • Figure 7: Final score per model, mean with standard deviation over the three seeds, individual seeds shown as points. Dashed and dotted lines mark the oracle and heuristic means. Stability separates models that a single-seed board cannot tell apart, from kimi-k2.6 at 0.15 to claude-haiku-4.5 at 22.73.
  • Figure 8: Transfer offers per season, solo track (rows sorted by final score; totals at right). The winner’s 9 offers are all in the first two seasons; the field-wide offer-to-completion conversion is 2–5%.
  • Figure 9: Transfer offers per season in the Arena (approximate year mapping; daggered models settled early). The winner’s offer count rises from 9 (solo) to 91: the same model under a different, correct judgment of market liquidity.
  • Figure 10: Season-by-season cumulative API cost vs. composite score, solo track. Steep curves convert dollars into score throughout; flat long curves do not. The full 20-year runs span $18 to $191.
  • Figure 11: The failed metric: share of stops carrying an engine warning, by type (left) and against final score (right), means over the three seeds. The per-seed coefficient runs +0.45 / +0.07 / −0.78 and averages to nothing, so warning exposure is not reported as a capability. Zero warnings conflates genuine anticipation with do-nothing conservatism.
  • Figure 12: Both panels are means over the three seeds. Left: youth harvest rate (share of promotions later fielded in ≥5 lineups); the correlation that looked informative on seed 1 does not survive three (−0.52 / −0.18 / +0.04). Right: endgame shift in long-horizon actions per season, years 2–16 (blue) vs. 17–20 (red); endgame-aware models reduce late investment, one ramps up.
  • Figure 13: Memory-curation regimes on the solo track. Left: season-to-season TF–IDF cosine similarity of each model’s reconstructed season-end notebook. Right: notebook size over the run. The right panel decodes the left: the same “high consistency” is a 200k-character append-only archive for one model and a 3–6k curated document for the winner.
  • Figure 14: Compute allocation. Mean tokens per run against mean solo score, both over the three seeds, on a log axis. A null result (rs=−0.19, p=0.50) across a sevenfold spend range.
  • Figure 15: The asset channels behind Figure 6: per-season net worth and squad value for all 16 Arena seats, the six highest-scoring seats colored as in Figure 6 and the grey lines the remaining ten (numeric snapshots in Table 9). The winner’s mid-table seasons are asset accumulation (the net-worth curve keeps climbing while the standings do not), and the conservative cash-holders’ net worth grows on cash the composite discounts.
在原文中查看图表 →

论文原文摘要(英文)

Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget a

作者 · Tianyou Wang, Chongyang Gao, Kezhen Chen, Chen Dong, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道