매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

arXiv:2608.184232026-08-20

AI에게 축구 구단을 20년간 통째로 맡겨보니, 승패는 계산력이 아니라 '경영 감각'에서 갈렸다

FM-Bench는 언어모델 에이전트가 26개 도구를 써서 20년(약 340~400번의 의사결정 순간) 동안 가상의 축구 구단을 운영하도록 만든 벤치마크다. 숨겨진 선수 능력치, 시간이 지나야 드러나는 투자 성과, 상대 반응에 따라 변하는 이적 시장, 성적과 재정을 동시에 평가하는 이사회라는 네 가지 장치로 '장기 경영'이라는 능력을 측정한다. 15개 최신 모델을 혼자 하는 모드와 16개 좌석이 한 세계를 공유하는 아레나 모드에 넣어 세 번씩 반복 실행한 결과, claude-fable-5가 두 트랙 모두에서 1위를 차지했지만 순위는 모델 크기나 가격, 개발사와 무관하게 갈렸다.

무엇을 했나

  1. 20년간 이어지는 가상 축구 구단 경영을 통해 언어모델이 장기간 누적되는 결정을 잘 내리는지 측정하는 벤치마크를 만들었다.
  2. 선수 능력치는 영원히 정확히 알 수 없고, 투자 성과는 몇 년 뒤에 나타나며, 이적 시장은 거절당한 제안에 반응해 가격을 올리고, 이사회는 성적과 재정을 함께 평가하는 방식으로 '경영'의 어려움을 구현했다.
  3. 15개 최신 모델(Anthropic, OpenAI, Google, xAI, Meta의 폐쇄형 모델과 5개 오픈웨이트 모델)을 혼자 뛰는 솔로 트랙과 15개 모델이 한 세계에서 경쟁하는 아레나 트랙에 각각 투입했다.
  4. 최종 점수를 결승선 감각으로 투자 줄이기, 현금을 놀리지 않기, 계약 갱신을 미리 여는지 등 6가지 행동 지표로 쪼개어 어떤 습관이 좋은 성적과 연결되는지 분석했다.
  5. claude-fable-5가 두 트랙에서 모두 1위(솔로 평균 90.94점, 참값을 아는 가상의 최고 기준점은 95.54점)를 기록했지만, 아레나에서는 우승 팀이 10개 모델 사이를 오갈 만큼 경쟁이 순위를 흔들었다.
Table 1: The four demands: the mechanism that instantiates each, the ability it targets, and the calibration lever that tunes it.
demandmechanismtargeted abilitycalibration lever
Hidden informationscout bands with a permanent per-scout bias; hidden player traits (injury proneness, development rate, aging onset); hidden asks in negotiationvaluation calibration under noisebandwidth, bias sigma
Cumulative consequencesUpside: youth development and facility investment pay off over years. Downside: insolvency ends in administration (points penalty, forced sales), confidence collapse ends in firing; both compound yet remain recoverable. Honors accrue into the final score season by seasonlong-range credit assignment; trend recognition, loss-cuttinggrowth and aging curves; spiral thresholds, warning cadence; honors weights
Counter-adaptive marketrejected bids raise the hidden ask; repeat-pair markups (bargains included); anti-inversion counters; negotiation cooldowns; per-seed mispricingstrategy adaptationmarkup and decay rates
Multi-objective pressurethe board judges results and financial discipline jointly; season targets scale with squad strengthmulti-constraint balancingtarget cushion, board patience
Table 2: The Solo track results. The four non-LLM seats are reference policies. The oracle reads true hidden state through the same interface (§4.1, Appendix 13); heuristic is a disciplined blind script, idle never acts, and random acts arbitrarily (Appendix 10).
seatSfinaltokens
oracle (privileged)95.54±4.68
claude-fable-590.94±5.2024M
kimi-k2.688.49±0.1587M
gpt-5.6-terra86.66±1.2028M
gpt-5.6-sol86.40±2.5362M
muse-spark-1.183.19±12.0373M
glm-5.283.17±0.9258M
grok-4.581.82±9.87191M
qwen3.7-max80.66±10.3847M
deepseek-v4-pro79.15±2.67194M
gemini-3-flash79.06±13.4551M
claude-sonnet-575.75±5.0339M
claude-opus-4.875.02±2.4631M
gemini-3.5-flash74.59±11.02154M
minimax-m368.37±12.3374M
claude-haiku-4.536.90±22.7386M
heuristic17.05±12.34
idle−0.90±1.86
random−17.21±2.45
Table 3: The six first-play human runs (§4.2.2). Deaths are capped revivals spent before the settling firing; t is the game-year at which a fired run settled.
playerSfinalsettledeathsstopstime
H174.64completed039410 h
H259.95completed03984 h
H35.39fired, t=10.842024 h
H42.47fired, t=10.741985 h
H50.41fired, t=7.841651.5 h
H6−3.78fired, t=12.342372 h
Table 4: The Arena results.
#seatSfinaldeathssettletokens
1claude-fable-576.260completed∼34M
2muse-spark-1.162.470completed∼80M
3deepseek-v4-pro52.090completed∼111M
4glm-5.251.440completed∼86M
5grok-4.550.890completed∼138M
6gpt-5.6-sol47.940completed∼78M
7minimax-m344.780completed∼65M
8claude-sonnet-540.750completed∼42M
9kimi-k2.639.720completed∼119M
10gpt-5.6-terra38.440completed∼19M
11qwen3.7-max32.772completed∼54M
12claude-opus-4.828.441completed∼43M
13gemini-3-flash15.762completed∼63M
14claude-haiku-4.50.764fired, t=10.7∼30M
15gemini-3.5-flash0.214fired, t=5.8∼40M
16heuristic (anchor)0.134fired, t=2.6
Table 5: The roster: 15 frontier LLM seats plus one scripted anchor, July 2026.
#seatproviderrole
1claude-fable-5 (4)Anthropicnew flagship
2claude-opus-4.8 (5)Anthropicprevious flagship
3claude-haiku-4.5 (3)Anthropicsmall tier
4gpt-5.6-sol (32)OpenAInew flagship
5gpt-5.6-terra (32)OpenAIbalanced tier
6claude-sonnet-5 (6)Anthropicworkhorse mid (4-point family curve)
7gemini-3.5-flash (16)Googlelatest GA flagship
8gemini-3-flash (15)Googleprevious-gen fast
9deepseek-v4-pro (40)Togetheropen flagship
10grok-4.5 (38)xAIfrontier
11kimi-k2.6 (30)Togetheropen agentic specialist
12qwen3.7-max (1)Togetheropen flagship
13glm-5.2 (42)Togetheropen top tier
14muse-spark-1.1 (27)Meta Model APIMeta agentic model
15minimax-m3 (29)Togetheropen flagship
16heuristic— (scripted)disciplined-script anchor
Table 6: The five competence levers. Medium is the neutral setting.
levereasymediumhard
scouting noise, multiplier on the belief sigma1.31.00.35
lineup rotates for fatiguenoyesyes
market days acted onevery 2ndeveryevery
gap required before buying an upgrade1.51.00.25
wage budget, multiplier on the board target0.921.01.10
Table 7: Per-model profiles over the three seeds. Score is the mean and SD the standard deviation of Sfinal; axis extremes come from the matrix of Figure 5.
seatscoresdprofile
claude-fable-590.95.2generalist, no axis below 0.79; best bidder in the field at 9 offers per signing and the sharpest endgame reduction
kimi-k2.688.50.1the steadiest model in the campaign, scoring within a third of a point on three different worlds
gpt-5.6-terra86.71.2minimalist: lowest token spend in the field (28M) after the winner and an early endgame reduction, but the shortest renewal leads among the top models
gpt-5.6-sol86.42.5archivist memory (similarity 0.91, append-only) and the weakest bidding of the top group (45 offers per signing)
muse-spark-1.183.212.0strong on two worlds and 20 points weaker on the third; one of two models that raise long-horizon spending at the end
glm-5.283.20.9stable across seeds but loose with cash (135% idle ratio)
grok-4.581.89.9compute as substitute: long renewal leads and heavy spend (191M tokens) with wide seed-to-seed swings
qwen3.7-max80.710.4longest renewal leads after the winner, undone by a 103% idle ratio and churning memory (similarity 0.23)
deepseek-v4-pro79.22.7near-static notebook (0.81) and the heaviest spend in the field (194M tokens)
gemini-3-flash79.113.5best cash discipline in the field (46%) but the least willing to cut endgame spending; collapses on the hardest seed
claude-sonnet-575.75.0churning memory (0.20, the lowest) and no endgame reduction
claude-opus-4.875.02.5shortest renewal leads in the field (10 months) with a 134% idle ratio
gemini-3.5-flash74.611.0worst price discovery (73 offers per signing) and the largest endgame ramp-up
minimax-m368.412.3prefers veteran signings, short leads, and swings 23 points across seeds
claude-haiku-4.536.922.7weak on every axis at once: 196% idle cash, 11-month leads, a rising endgame rate, and the widest spread in the campaign
Table 8: Solo season snapshots (seed 1). Cell format: position ⋅ honors ⋅ net worth ⋅ squad value ⋅ score. Rows sorted by final score.
seatY5Y10Y15Y20
claude-fable-54th ⋅ 140H ⋅ 293M ⋅ 130M ⋅ 33.31st ⋅ 480H ⋅ 470M ⋅ 218M ⋅ 50.31st ⋅ 900H ⋅ 1048M ⋅ 443M ⋅ 65.91st ⋅ 1400H ⋅ 1742M ⋅ 634M ⋅ 94.6
muse-spark-1.13rd ⋅ 120H ⋅ 302M ⋅ 158M ⋅ 31.11st ⋅ 460H ⋅ 468M ⋅ 196M ⋅ 48.71st ⋅ 960H ⋅ 1037M ⋅ 457M ⋅ 66.51st ⋅ 1460H ⋅ 1298M ⋅ 315M ⋅ 91.2
grok-4.54th ⋅ 60H ⋅ 289M ⋅ 169M ⋅ 23.91st ⋅ 460H ⋅ 318M ⋅ 107M ⋅ 46.61st ⋅ 880H ⋅ 618M ⋅ 225M ⋅ 58.11st ⋅ 1380H ⋅ 1106M ⋅ 634M ⋅ 89.7
kimi-k2.63rd ⋅ 120H ⋅ 344M ⋅ 116M ⋅ 31.91st ⋅ 440H ⋅ 430M ⋅ 99M ⋅ 46.71st ⋅ 940H ⋅ 668M ⋅ 276M ⋅ 61.41st ⋅ 1440H ⋅ 733M ⋅ 400M ⋅ 88.5
gemini-3-flash3rd ⋅ 160H ⋅ 247M ⋅ 165M ⋅ 32.81st ⋅ 500H ⋅ 339M ⋅ 111M ⋅ 42.11st ⋅ 1000H ⋅ 675M ⋅ 400M ⋅ 61.51st ⋅ 1420H ⋅ 812M ⋅ 618M ⋅ 86.6
qwen3.7-max3rd ⋅ 120H ⋅ 404M ⋅ 127M ⋅ 34.01st ⋅ 360H ⋅ 582M ⋅ 140M ⋅ 48.01st ⋅ 780H ⋅ 917M ⋅ 256M ⋅ 62.31st ⋅ 1200H ⋅ 1305M ⋅ 248M ⋅ 86.1
gpt-5.6-terra3rd ⋅ 140H ⋅ 359M ⋅ 120M ⋅ 34.01st ⋅ 440H ⋅ 565M ⋅ 102M ⋅ 50.11st ⋅ 780H ⋅ 808M ⋅ 247M ⋅ 59.81st ⋅ 1280H ⋅ 987M ⋅ 186M ⋅ 85.3
gpt-5.6-sol4th ⋅ 140H ⋅ 343M ⋅ 135M ⋅ 33.61st ⋅ 460H ⋅ 615M ⋅ 176M ⋅ 51.51st ⋅ 960H ⋅ 762M ⋅ 217M ⋅ 62.62nd ⋅ 1220H ⋅ 1109M ⋅ 274M ⋅ 84.5
glm-5.23rd ⋅ 120H ⋅ 294M ⋅ 121M ⋅ 31.81st ⋅ 420H ⋅ 513M ⋅ 110M ⋅ 48.13rd ⋅ 760H ⋅ 712M ⋅ 58M ⋅ 56.61st ⋅ 1080H ⋅ 1331M ⋅ 266M ⋅ 82.8
minimax-m310th ⋅ 20H ⋅ 324M ⋅ 154M ⋅ 20.91st ⋅ 240H ⋅ 492M ⋅ 100M ⋅ 39.31st ⋅ 540H ⋅ 827M ⋅ 155M ⋅ 51.81st ⋅ 1040H ⋅ 1245M ⋅ 327M ⋅ 82.5
claude-sonnet-55th ⋅ 20H ⋅ 247M ⋅ 91M ⋅ 18.62nd ⋅ 180H ⋅ 363M ⋅ 97M ⋅ 36.01st ⋅ 480H ⋅ 647M ⋅ 94M ⋅ 50.21st ⋅ 980H ⋅ 1020M ⋅ 146M ⋅ 80.4
deepseek-v4-pro5th ⋅ 140H ⋅ 385M ⋅ 124M ⋅ 35.81st ⋅ 480H ⋅ 638M ⋅ 111M ⋅ 52.91st ⋅ 820H ⋅ 732M ⋅ 110M ⋅ 58.73rd ⋅ 1080H ⋅ 996M ⋅ 163M ⋅ 79.0
claude-opus-4.83rd ⋅ 140H ⋅ 308M ⋅ 100M ⋅ 32.71st ⋅ 280H ⋅ 467M ⋅ 78M ⋅ 40.52nd ⋅ 600H ⋅ 788M ⋅ 160M ⋅ 56.43rd ⋅ 760H ⋅ 1313M ⋅ 216M ⋅ 76.9
gemini-3.5-flash2nd ⋅ 140H ⋅ 260M ⋅ 186M ⋅ 29.91st ⋅ 480H ⋅ 226M ⋅ 140M ⋅ 41.31st ⋅ 980H ⋅ 479M ⋅ 156M ⋅ 57.25th ⋅ 1040H ⋅ 877M ⋅ 170M ⋅ 75.2
claude-haiku-4.512th ⋅ 100H ⋅ 293M ⋅ 132M ⋅ 28.47th ⋅ 200H ⋅ 395M ⋅ 39M ⋅ 33.83rd ⋅ 340H ⋅ 629M ⋅ 42M ⋅ 42.812th ⋅ 360H ⋅ 821M ⋅ 59M ⋅ 57.3
Table 9: Arena season snapshots (seed 7, one shared world); the 15 models, as on the public page (the scripted anchor settled at t≈2.6y). Cell format as in Table 8. Seats that exhausted the revival cap keep their frozen score afterward, visible as repeated values (claude-haiku-4.5 from Y15, gemini-3.5-flash from Y10). Rows sorted by final score.
seatY5Y10Y15Y20
claude-fable-52nd ⋅ 180H ⋅ 326M ⋅ 105M ⋅ 47.45th ⋅ 240H ⋅ 394M ⋅ 196M ⋅ 56.45th ⋅ 440H ⋅ 568M ⋅ 250M ⋅ 71.33rd ⋅ 620H ⋅ 507M ⋅ 298M ⋅ 76.3
muse-spark-1.18th ⋅ 0H ⋅ 260M ⋅ 84M ⋅ 18.57th ⋅ 40H ⋅ 197M ⋅ 117M ⋅ 24.02nd ⋅ 160H ⋅ 272M ⋅ 222M ⋅ 46.21st ⋅ 420H ⋅ 306M ⋅ 246M ⋅ 62.5
deepseek-v4-pro9th ⋅ 0H ⋅ 279M ⋅ 91M ⋅ 20.010th ⋅ 140H ⋅ 342M ⋅ 91M ⋅ 44.54th ⋅ 200H ⋅ 397M ⋅ 122M ⋅ 52.25th ⋅ 200H ⋅ 363M ⋅ 160M ⋅ 52.1
glm-5.24th ⋅ 40H ⋅ 275M ⋅ 107M ⋅ 30.31st ⋅ 160H ⋅ 335M ⋅ 78M ⋅ 46.18th ⋅ 200H ⋅ 346M ⋅ 71M ⋅ 49.36th ⋅ 200H ⋅ 356M ⋅ 125M ⋅ 51.4
grok-4.57th ⋅ 20H ⋅ 197M ⋅ 134M ⋅ 21.16th ⋅ 260H ⋅ 145M ⋅ 145M ⋅ 40.07th ⋅ 320H ⋅ 108M ⋅ 185M ⋅ 36.48th ⋅ 520H ⋅ 139M ⋅ 199M ⋅ 50.9
gpt-5.6-sol6th ⋅ 100H ⋅ 258M ⋅ 115M ⋅ 36.94th ⋅ 120H ⋅ 239M ⋅ 86M ⋅ 36.81st ⋅ 240H ⋅ 281M ⋅ 229M ⋅ 52.32nd ⋅ 280H ⋅ 182M ⋅ 294M ⋅ 47.9
minimax-m33rd ⋅ 120H ⋅ 272M ⋅ 93M ⋅ 38.211th ⋅ 220H ⋅ 246M ⋅ 61M ⋅ 43.213th ⋅ 220H ⋅ 236M ⋅ 64M ⋅ 42.513th ⋅ 220H ⋅ 260M ⋅ 76M ⋅ 44.8
claude-sonnet-511th ⋅ 100H ⋅ 246M ⋅ 82M ⋅ 35.913th ⋅ 100H ⋅ 268M ⋅ 64M ⋅ 36.49th ⋅ 100H ⋅ 251M ⋅ 77M ⋅ 36.010th ⋅ 160H ⋅ 226M ⋅ 94M ⋅ 40.8
kimi-k2.613th ⋅ 0H ⋅ 238M ⋅ 76M ⋅ 18.18th ⋅ 20H ⋅ 305M ⋅ 82M ⋅ 26.93rd ⋅ 60H ⋅ 351M ⋅ 106M ⋅ 36.97th ⋅ 80H ⋅ 304M ⋅ 175M ⋅ 39.7
gpt-5.6-terra5th ⋅ 40H ⋅ 266M ⋅ 108M ⋅ 28.82nd ⋅ 60H ⋅ 113M ⋅ 108M ⋅ 13.06th ⋅ 60H ⋅ 352M ⋅ 108M ⋅ 36.29th ⋅ 60H ⋅ 390M ⋅ 138M ⋅ 38.4
qwen3.7-max14th ⋅ 20H ⋅ 279M ⋅ 87M ⋅ 19.212th ⋅ 20H ⋅ 254M ⋅ 88M ⋅ 13.612th ⋅ 180H ⋅ 451M ⋅ 200M ⋅ 34.211th ⋅ 180H ⋅ 451M ⋅ 113M ⋅ 32.8
claude-opus-4.810th ⋅ 20H ⋅ 274M ⋅ 90M ⋅ 18.43rd ⋅ 40H ⋅ 263M ⋅ 99M ⋅ 21.210th ⋅ 40H ⋅ 412M ⋅ 132M ⋅ 27.814th ⋅ 40H ⋅ 472M ⋅ 95M ⋅ 28.4
gemini-3-flash1st ⋅ 140H ⋅ −6M ⋅ 110M ⋅ 12.79th ⋅ 160H ⋅ −77M ⋅ 107M ⋅ 10.711th ⋅ 160H ⋅ 92M ⋅ 111M ⋅ 11.74th ⋅ 200H ⋅ 107M ⋅ 182M ⋅ 15.8
claude-haiku-4.516th ⋅ 0H ⋅ 209M ⋅ 64M ⋅ 11.014th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 4.314th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 0.816th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 0.8
gemini-3.5-flash12th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 2.115th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.215th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.212th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.2
Table 10: The 26-tool interface: 12 query tools, 11 action tools, 2 notebook tools, and 1 control tool. Descriptions follow the tool schemas served to the agent.
toolwhat it does
Query (12), charged to the 30-query budget:
get_club_overviewthe agent’s club at a glance: division, cash, board, facilities, tactics
get_squadthe senior squad with ability bands, form, contracts, wages
get_playerdetailed view of any player by id
get_league_tableleague standings
get_fixturesthe club’s season fixtures and results
get_match_reportmatch report by match id
get_transfer_marketplayers available to buy: listed, free agents, expiring contracts
get_financescash, revenue, wage bill, active warnings
get_youth_academyacademy prospects with potential stars
get_historythe run’s archive: honors, seasons, transfers
get_inboxpending offers and open items
get_draft_pooldraft phase only: the shared template pool with scouted ability bands, potential stars, prices, and preset contracts, plus budget status
Action (11); negotiation moves charged to the 10-move budget:
set_lineupset the preferred starting XI (auto-completed if players are unavailable)
set_tacticsset formation and playing style from fixed enums
make_transfer_offerbid for another club’s player; the negotiation resolves within the stop
respond_to_offeraccept, reject, counter, or accept a counter on a pending offer
offer_contractrenew an own player or sign a free agent at a wage and term
list_playerput a player on or off the transfer list
release_playerterminate a contract (severance: half of one year’s wage)
promote_youthpromote an academy prospect to the senior squad
investstart a facility upgrade: academy, training, or stadium
set_standing_orderadd or clear a standing order (max 10; one threshold plus a fixed action)
submit_draftdraft phase only: submit the complete pick list; resubmitting replaces it, and auto-fill completes a short list from the cheapest tier
Notebook (2), free:
append_noteappend to the private notebook (persists across stops)
rewrite_notesreplace the entire notebook
Control (1):
advanceend this decision stop and simulate to the next one
Table 11: The decision-stop calendar: 13 scheduled stops per season on the 360-day year, plus event-driven stops the world raises unsolicited. Several types can fire at one stop.
stop (day of the 360-day year)what the world wants decided
Scheduled, 13 per season:
Preseason board (2)the board sets the season target; plan the year
Summer window plan (5)the transfer window opens; build the squad
Summer deadline (59)last actions before the window closes
Lineup lock (63)commit lineup and tactics before round 1
Monthly report (90, 120, 150, 240)finances and form check-in
Winter window open (181)mid-course squad correction
Midseason review (190)the board measures progress against target
Winter deadline (209)last winter-window actions
Season settlement (300)the season closes: honors, revenue, board verdict
Youth intake (305)the academy class arrives; promote, hold, or release
Event-driven, unsolicited:
Draft (once, at run start)assemble the 25-player squad from the shared pool
Offer batchincoming bids for the agent’s players, batched at a minimum 15-day spacing
Contract expirya senior contract is running down; renew or lose the player for free
Major injurya first-team player goes down; cover or reshuffle
Board warningconfidence is deteriorating; the board expects a response
Insolvency warningcash runway is shrinking (30- and 60-day warnings)
Administrationinsolvency executed: points penalty and forced sales
Revivalthe seat restarts under the capped-revival mechanism (§4.2)
Table 12: The four memory layers (§3.2). What happened is recorded automatically; what it means and what to do next must be written by the agent itself.
layerwho writes itcontentsdelivery
Packet auto-contextenginedashboard digest, inbox (all events since last stop), agent’s last 10 engine-mutating actions, pending offerspushed free into every stop packet
Archiveenginehonors / season / transfer history of the whole runget_history query tool, on demand
Notebookthe agentintentions, reasoning, plans (“why I bought him”, “sell Y in winter”)injected into every packet; written via append_note / rewrite_notes
Audit logrunnerevery successful actionnever fed back: replay/anti-cheat/resume only
Table 13: One 20-year run, by the numbers: medians and ranges across the 15 solo models.
quantitymedianrange
decision stops374341–390
of which scheduled∼260 (13/season)rest event-triggered (Table 11)
matches simulated in the world4,800(30 rounds × 8 matches × 20 seasons)
model turns (API calls)∼1,7601,184–6,969
state-changing actions849–2,198
tokens (manifest accounting)61M24M–274M
notebook writes354320–425
notebook size at year 200.4k–209k characters
simulation parameters (frozen)313

왜 중요한가

짧은 문제 하나를 잘 푸는 것과 수년에 걸쳐 누적되는 결정을 잘 관리하는 것은 전혀 다른 능력이라는 점을 실제로 측정 가능하게 만들었다는 데 의미가 있다. 앞으로 에이전트를 회사 운영, 장기 투자, 자원 관리처럼 실제 세계의 '경영' 업무에 투입하려는 사람들에게, 모델 크기나 가격표가 아니라 어떤 행동 습관을 봐야 하는지에 대한 구체적인 기준을 제시한다.

이 논문의 용어

  • 에이전트(agent) · 스스로 도구를 골라 실행하며 목표를 수행하는 AI 시스템
  • 솔로 트랙 / 아레나 · 솔로 트랙은 모델 혼자 고정된 상대와 겨루는 모드, 아레나는 여러 모델이 하나의 세계를 공유하며 직접 경쟁하는 모드
  • 오라클(oracle) · 숨겨진 정보를 모두 볼 수 있는 특권을 가진 비교용 스크립트, 최고 성능의 기준선 역할
  • 노트북(notebook) · 매 순간 대화가 초기화되는 환경에서 에이전트가 스스로 남기는 유일한 기록, 기억 유지 전략 자체가 평가 대상
  • Sfinal(최종 점수) · 우승 등 성과 점수, 순자산 증가분, 스쿼드 가치를 합쳐 20년 전체를 누적 평가하는 단일 점수

본문에 싣지 못한 그림

  • Figure 1: How a run works. One run spans 20 in-game years and roughly 340 to 400 decision stops, ending in a single composite score. At each stop, the clock freezes and the agent takes as many tool-call turns as it wants (read → act → advance). The agent’s only cross-stop memory is its self-written notebook (§3.2); the environment side holds the club (net worth, squad with permanently hidden true ability, a board that can fire the manager) and the world (15 rival clubs, a counter-adaptive market, delayed-payoff investments).
  • Figure 2: Per-season solo trajectories for all 15 models in four channels. Lines are the mean over three seeds, shading spans the best and worst seed, and the five models with the highest mean are colored and named in the legend; the grey lines are the remaining ten models.
  • Figure 3: Credit assignment. Left: discretionary spend bucketed by payoff horizon, with run totals at right. Right: mean season-end idle-cash ratio vs. final score, rs=−0.50, negative on every seed.
  • Figure 4: Proactive control. Left: distribution of contract months remaining when each renewal negotiation episode opens (rows sorted by final score; shaded band = last-minute, ≤6 months). Right: per-model median lead vs. final score, rs=+0.45, positive on every seed.
  • Figure 5: Capability matrix over the 15 solo models, six axes. Each column rank-normalizes one behavioral metric averaged over the three seeds (1 = best of 15; construction in Appendix 15), and rows are sorted by mean final score. The winner is a generalist, with no axis below 0.79 and mean 0.94, rather than the leader of any single one; the mid-table pairs a genuine strength with a decisive gap; and the bottom of the board is weak on most axes at once.
  • Figure 6: Arena trajectories, all 16 seats. The six highest-scoring seats are colored and named in the legend; the grey lines are the remaining ten seats. Left: composite Sfinal at each season end. Right: league-table position.
  • Figure 7: Final score per model, mean with standard deviation over the three seeds, individual seeds shown as points. Dashed and dotted lines mark the oracle and heuristic means. Stability separates models that a single-seed board cannot tell apart, from kimi-k2.6 at 0.15 to claude-haiku-4.5 at 22.73.
  • Figure 8: Transfer offers per season, solo track (rows sorted by final score; totals at right). The winner’s 9 offers are all in the first two seasons; the field-wide offer-to-completion conversion is 2–5%.
  • Figure 9: Transfer offers per season in the Arena (approximate year mapping; daggered models settled early). The winner’s offer count rises from 9 (solo) to 91: the same model under a different, correct judgment of market liquidity.
  • Figure 10: Season-by-season cumulative API cost vs. composite score, solo track. Steep curves convert dollars into score throughout; flat long curves do not. The full 20-year runs span $18 to $191.
  • Figure 11: The failed metric: share of stops carrying an engine warning, by type (left) and against final score (right), means over the three seeds. The per-seed coefficient runs +0.45 / +0.07 / −0.78 and averages to nothing, so warning exposure is not reported as a capability. Zero warnings conflates genuine anticipation with do-nothing conservatism.
  • Figure 12: Both panels are means over the three seeds. Left: youth harvest rate (share of promotions later fielded in ≥5 lineups); the correlation that looked informative on seed 1 does not survive three (−0.52 / −0.18 / +0.04). Right: endgame shift in long-horizon actions per season, years 2–16 (blue) vs. 17–20 (red); endgame-aware models reduce late investment, one ramps up.
  • Figure 13: Memory-curation regimes on the solo track. Left: season-to-season TF–IDF cosine similarity of each model’s reconstructed season-end notebook. Right: notebook size over the run. The right panel decodes the left: the same “high consistency” is a 200k-character append-only archive for one model and a 3–6k curated document for the winner.
  • Figure 14: Compute allocation. Mean tokens per run against mean solo score, both over the three seeds, on a log axis. A null result (rs=−0.19, p=0.50) across a sevenfold spend range.
  • Figure 15: The asset channels behind Figure 6: per-season net worth and squad value for all 16 Arena seats, the six highest-scoring seats colored as in Figure 6 and the grey lines the remaining ten (numeric snapshots in Table 9). The winner’s mid-table seasons are asset accumulation (the net-worth curve keeps climbing while the standings do not), and the conservative cash-holders’ net worth grows on cash the composite discounts.
원문에서 그림 보기 →

논문 원문 초록 (영문)

Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget a

저자 · Tianyou Wang, Chongyang Gao, Kezhen Chen, Chen Dong, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사