FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
AI에게 축구 구단을 20년간 통째로 맡겨보니, 승패는 계산력이 아니라 '경영 감각'에서 갈렸다
FM-Bench는 언어모델 에이전트가 26개 도구를 써서 20년(약 340~400번의 의사결정 순간) 동안 가상의 축구 구단을 운영하도록 만든 벤치마크다. 숨겨진 선수 능력치, 시간이 지나야 드러나는 투자 성과, 상대 반응에 따라 변하는 이적 시장, 성적과 재정을 동시에 평가하는 이사회라는 네 가지 장치로 '장기 경영'이라는 능력을 측정한다. 15개 최신 모델을 혼자 하는 모드와 16개 좌석이 한 세계를 공유하는 아레나 모드에 넣어 세 번씩 반복 실행한 결과, claude-fable-5가 두 트랙 모두에서 1위를 차지했지만 순위는 모델 크기나 가격, 개발사와 무관하게 갈렸다.
무엇을 했나
- 20년간 이어지는 가상 축구 구단 경영을 통해 언어모델이 장기간 누적되는 결정을 잘 내리는지 측정하는 벤치마크를 만들었다.
- 선수 능력치는 영원히 정확히 알 수 없고, 투자 성과는 몇 년 뒤에 나타나며, 이적 시장은 거절당한 제안에 반응해 가격을 올리고, 이사회는 성적과 재정을 함께 평가하는 방식으로 '경영'의 어려움을 구현했다.
- 15개 최신 모델(Anthropic, OpenAI, Google, xAI, Meta의 폐쇄형 모델과 5개 오픈웨이트 모델)을 혼자 뛰는 솔로 트랙과 15개 모델이 한 세계에서 경쟁하는 아레나 트랙에 각각 투입했다.
- 최종 점수를 결승선 감각으로 투자 줄이기, 현금을 놀리지 않기, 계약 갱신을 미리 여는지 등 6가지 행동 지표로 쪼개어 어떤 습관이 좋은 성적과 연결되는지 분석했다.
- claude-fable-5가 두 트랙에서 모두 1위(솔로 평균 90.94점, 참값을 아는 가상의 최고 기준점은 95.54점)를 기록했지만, 아레나에서는 우승 팀이 10개 모델 사이를 오갈 만큼 경쟁이 순위를 흔들었다.
| demand | mechanism | targeted ability | calibration lever |
|---|---|---|---|
| Hidden information | scout bands with a permanent per-scout bias; hidden player traits (injury proneness, development rate, aging onset); hidden asks in negotiation | valuation calibration under noise | bandwidth, bias sigma |
| Cumulative consequences | Upside: youth development and facility investment pay off over years. Downside: insolvency ends in administration (points penalty, forced sales), confidence collapse ends in firing; both compound yet remain recoverable. Honors accrue into the final score season by season | long-range credit assignment; trend recognition, loss-cutting | growth and aging curves; spiral thresholds, warning cadence; honors weights |
| Counter-adaptive market | rejected bids raise the hidden ask; repeat-pair markups (bargains included); anti-inversion counters; negotiation cooldowns; per-seed mispricing | strategy adaptation | markup and decay rates |
| Multi-objective pressure | the board judges results and financial discipline jointly; season targets scale with squad strength | multi-constraint balancing | target cushion, board patience |
| seat | Sfinal | tokens |
|---|---|---|
| oracle (privileged) | 95.54±4.68 | — |
| claude-fable-5 | 90.94±5.20 | 24M |
| kimi-k2.6 | 88.49±0.15 | 87M |
| gpt-5.6-terra | 86.66±1.20 | 28M |
| gpt-5.6-sol | 86.40±2.53 | 62M |
| muse-spark-1.1 | 83.19±12.03 | 73M |
| glm-5.2 | 83.17±0.92 | 58M |
| grok-4.5 | 81.82±9.87 | 191M |
| qwen3.7-max | 80.66±10.38 | 47M |
| deepseek-v4-pro | 79.15±2.67 | 194M |
| gemini-3-flash | 79.06±13.45 | 51M |
| claude-sonnet-5 | 75.75±5.03 | 39M |
| claude-opus-4.8 | 75.02±2.46 | 31M |
| gemini-3.5-flash | 74.59±11.02 | 154M |
| minimax-m3 | 68.37±12.33 | 74M |
| claude-haiku-4.5 | 36.90±22.73 | 86M |
| heuristic | 17.05±12.34 | — |
| idle | −0.90±1.86 | — |
| random | −17.21±2.45 | — |
| player | Sfinal | settle | deaths | stops | time |
|---|---|---|---|---|---|
| H1 | 74.64 | completed | 0 | 394 | 10 h |
| H2 | 59.95 | completed | 0 | 398 | 4 h |
| H3 | 5.39 | fired, t=10.8 | 4 | 202 | 4 h |
| H4 | 2.47 | fired, t=10.7 | 4 | 198 | 5 h |
| H5 | 0.41 | fired, t=7.8 | 4 | 165 | 1.5 h |
| H6 | −3.78 | fired, t=12.3 | 4 | 237 | 2 h |
| # | seat | Sfinal | deaths | settle | tokens |
|---|---|---|---|---|---|
| 1 | claude-fable-5 | 76.26 | 0 | completed | ∼34M |
| 2 | muse-spark-1.1 | 62.47 | 0 | completed | ∼80M |
| 3 | deepseek-v4-pro | 52.09 | 0 | completed | ∼111M |
| 4 | glm-5.2 | 51.44 | 0 | completed | ∼86M |
| 5 | grok-4.5 | 50.89 | 0 | completed | ∼138M |
| 6 | gpt-5.6-sol | 47.94 | 0 | completed | ∼78M |
| 7 | minimax-m3 | 44.78 | 0 | completed | ∼65M |
| 8 | claude-sonnet-5 | 40.75 | 0 | completed | ∼42M |
| 9 | kimi-k2.6 | 39.72 | 0 | completed | ∼119M |
| 10 | gpt-5.6-terra | 38.44 | 0 | completed | ∼19M |
| 11 | qwen3.7-max | 32.77 | 2 | completed | ∼54M |
| 12 | claude-opus-4.8 | 28.44 | 1 | completed | ∼43M |
| 13 | gemini-3-flash | 15.76 | 2 | completed | ∼63M |
| 14 | claude-haiku-4.5 | 0.76 | 4 | fired, t=10.7 | ∼30M |
| 15 | gemini-3.5-flash | 0.21 | 4 | fired, t=5.8 | ∼40M |
| 16 | heuristic (anchor) | 0.13 | 4 | fired, t=2.6 | — |
| # | seat | provider | role |
|---|---|---|---|
| 1 | claude-fable-5 (4) | Anthropic | new flagship |
| 2 | claude-opus-4.8 (5) | Anthropic | previous flagship |
| 3 | claude-haiku-4.5 (3) | Anthropic | small tier |
| 4 | gpt-5.6-sol (32) | OpenAI | new flagship |
| 5 | gpt-5.6-terra (32) | OpenAI | balanced tier |
| 6 | claude-sonnet-5 (6) | Anthropic | workhorse mid (4-point family curve) |
| 7 | gemini-3.5-flash (16) | latest GA flagship | |
| 8 | gemini-3-flash (15) | previous-gen fast | |
| 9 | deepseek-v4-pro (40) | Together | open flagship |
| 10 | grok-4.5 (38) | xAI | frontier |
| 11 | kimi-k2.6 (30) | Together | open agentic specialist |
| 12 | qwen3.7-max (1) | Together | open flagship |
| 13 | glm-5.2 (42) | Together | open top tier |
| 14 | muse-spark-1.1 (27) | Meta Model API | Meta agentic model |
| 15 | minimax-m3 (29) | Together | open flagship |
| 16 | heuristic | — (scripted) | disciplined-script anchor |
| lever | easy | medium | hard |
|---|---|---|---|
| scouting noise, multiplier on the belief sigma | 1.3 | 1.0 | 0.35 |
| lineup rotates for fatigue | no | yes | yes |
| market days acted on | every 2nd | every | every |
| gap required before buying an upgrade | 1.5 | 1.0 | 0.25 |
| wage budget, multiplier on the board target | 0.92 | 1.0 | 1.10 |
| seat | score | sd | profile |
|---|---|---|---|
| claude-fable-5 | 90.9 | 5.2 | generalist, no axis below 0.79; best bidder in the field at 9 offers per signing and the sharpest endgame reduction |
| kimi-k2.6 | 88.5 | 0.1 | the steadiest model in the campaign, scoring within a third of a point on three different worlds |
| gpt-5.6-terra | 86.7 | 1.2 | minimalist: lowest token spend in the field (28M) after the winner and an early endgame reduction, but the shortest renewal leads among the top models |
| gpt-5.6-sol | 86.4 | 2.5 | archivist memory (similarity 0.91, append-only) and the weakest bidding of the top group (45 offers per signing) |
| muse-spark-1.1 | 83.2 | 12.0 | strong on two worlds and 20 points weaker on the third; one of two models that raise long-horizon spending at the end |
| glm-5.2 | 83.2 | 0.9 | stable across seeds but loose with cash (135% idle ratio) |
| grok-4.5 | 81.8 | 9.9 | compute as substitute: long renewal leads and heavy spend (191M tokens) with wide seed-to-seed swings |
| qwen3.7-max | 80.7 | 10.4 | longest renewal leads after the winner, undone by a 103% idle ratio and churning memory (similarity 0.23) |
| deepseek-v4-pro | 79.2 | 2.7 | near-static notebook (0.81) and the heaviest spend in the field (194M tokens) |
| gemini-3-flash | 79.1 | 13.5 | best cash discipline in the field (46%) but the least willing to cut endgame spending; collapses on the hardest seed |
| claude-sonnet-5 | 75.7 | 5.0 | churning memory (0.20, the lowest) and no endgame reduction |
| claude-opus-4.8 | 75.0 | 2.5 | shortest renewal leads in the field (10 months) with a 134% idle ratio |
| gemini-3.5-flash | 74.6 | 11.0 | worst price discovery (73 offers per signing) and the largest endgame ramp-up |
| minimax-m3 | 68.4 | 12.3 | prefers veteran signings, short leads, and swings 23 points across seeds |
| claude-haiku-4.5 | 36.9 | 22.7 | weak on every axis at once: 196% idle cash, 11-month leads, a rising endgame rate, and the widest spread in the campaign |
| seat | Y5 | Y10 | Y15 | Y20 |
|---|---|---|---|---|
| claude-fable-5 | 4th ⋅ 140H ⋅ 293M ⋅ 130M ⋅ 33.3 | 1st ⋅ 480H ⋅ 470M ⋅ 218M ⋅ 50.3 | 1st ⋅ 900H ⋅ 1048M ⋅ 443M ⋅ 65.9 | 1st ⋅ 1400H ⋅ 1742M ⋅ 634M ⋅ 94.6 |
| muse-spark-1.1 | 3rd ⋅ 120H ⋅ 302M ⋅ 158M ⋅ 31.1 | 1st ⋅ 460H ⋅ 468M ⋅ 196M ⋅ 48.7 | 1st ⋅ 960H ⋅ 1037M ⋅ 457M ⋅ 66.5 | 1st ⋅ 1460H ⋅ 1298M ⋅ 315M ⋅ 91.2 |
| grok-4.5 | 4th ⋅ 60H ⋅ 289M ⋅ 169M ⋅ 23.9 | 1st ⋅ 460H ⋅ 318M ⋅ 107M ⋅ 46.6 | 1st ⋅ 880H ⋅ 618M ⋅ 225M ⋅ 58.1 | 1st ⋅ 1380H ⋅ 1106M ⋅ 634M ⋅ 89.7 |
| kimi-k2.6 | 3rd ⋅ 120H ⋅ 344M ⋅ 116M ⋅ 31.9 | 1st ⋅ 440H ⋅ 430M ⋅ 99M ⋅ 46.7 | 1st ⋅ 940H ⋅ 668M ⋅ 276M ⋅ 61.4 | 1st ⋅ 1440H ⋅ 733M ⋅ 400M ⋅ 88.5 |
| gemini-3-flash | 3rd ⋅ 160H ⋅ 247M ⋅ 165M ⋅ 32.8 | 1st ⋅ 500H ⋅ 339M ⋅ 111M ⋅ 42.1 | 1st ⋅ 1000H ⋅ 675M ⋅ 400M ⋅ 61.5 | 1st ⋅ 1420H ⋅ 812M ⋅ 618M ⋅ 86.6 |
| qwen3.7-max | 3rd ⋅ 120H ⋅ 404M ⋅ 127M ⋅ 34.0 | 1st ⋅ 360H ⋅ 582M ⋅ 140M ⋅ 48.0 | 1st ⋅ 780H ⋅ 917M ⋅ 256M ⋅ 62.3 | 1st ⋅ 1200H ⋅ 1305M ⋅ 248M ⋅ 86.1 |
| gpt-5.6-terra | 3rd ⋅ 140H ⋅ 359M ⋅ 120M ⋅ 34.0 | 1st ⋅ 440H ⋅ 565M ⋅ 102M ⋅ 50.1 | 1st ⋅ 780H ⋅ 808M ⋅ 247M ⋅ 59.8 | 1st ⋅ 1280H ⋅ 987M ⋅ 186M ⋅ 85.3 |
| gpt-5.6-sol | 4th ⋅ 140H ⋅ 343M ⋅ 135M ⋅ 33.6 | 1st ⋅ 460H ⋅ 615M ⋅ 176M ⋅ 51.5 | 1st ⋅ 960H ⋅ 762M ⋅ 217M ⋅ 62.6 | 2nd ⋅ 1220H ⋅ 1109M ⋅ 274M ⋅ 84.5 |
| glm-5.2 | 3rd ⋅ 120H ⋅ 294M ⋅ 121M ⋅ 31.8 | 1st ⋅ 420H ⋅ 513M ⋅ 110M ⋅ 48.1 | 3rd ⋅ 760H ⋅ 712M ⋅ 58M ⋅ 56.6 | 1st ⋅ 1080H ⋅ 1331M ⋅ 266M ⋅ 82.8 |
| minimax-m3 | 10th ⋅ 20H ⋅ 324M ⋅ 154M ⋅ 20.9 | 1st ⋅ 240H ⋅ 492M ⋅ 100M ⋅ 39.3 | 1st ⋅ 540H ⋅ 827M ⋅ 155M ⋅ 51.8 | 1st ⋅ 1040H ⋅ 1245M ⋅ 327M ⋅ 82.5 |
| claude-sonnet-5 | 5th ⋅ 20H ⋅ 247M ⋅ 91M ⋅ 18.6 | 2nd ⋅ 180H ⋅ 363M ⋅ 97M ⋅ 36.0 | 1st ⋅ 480H ⋅ 647M ⋅ 94M ⋅ 50.2 | 1st ⋅ 980H ⋅ 1020M ⋅ 146M ⋅ 80.4 |
| deepseek-v4-pro | 5th ⋅ 140H ⋅ 385M ⋅ 124M ⋅ 35.8 | 1st ⋅ 480H ⋅ 638M ⋅ 111M ⋅ 52.9 | 1st ⋅ 820H ⋅ 732M ⋅ 110M ⋅ 58.7 | 3rd ⋅ 1080H ⋅ 996M ⋅ 163M ⋅ 79.0 |
| claude-opus-4.8 | 3rd ⋅ 140H ⋅ 308M ⋅ 100M ⋅ 32.7 | 1st ⋅ 280H ⋅ 467M ⋅ 78M ⋅ 40.5 | 2nd ⋅ 600H ⋅ 788M ⋅ 160M ⋅ 56.4 | 3rd ⋅ 760H ⋅ 1313M ⋅ 216M ⋅ 76.9 |
| gemini-3.5-flash | 2nd ⋅ 140H ⋅ 260M ⋅ 186M ⋅ 29.9 | 1st ⋅ 480H ⋅ 226M ⋅ 140M ⋅ 41.3 | 1st ⋅ 980H ⋅ 479M ⋅ 156M ⋅ 57.2 | 5th ⋅ 1040H ⋅ 877M ⋅ 170M ⋅ 75.2 |
| claude-haiku-4.5 | 12th ⋅ 100H ⋅ 293M ⋅ 132M ⋅ 28.4 | 7th ⋅ 200H ⋅ 395M ⋅ 39M ⋅ 33.8 | 3rd ⋅ 340H ⋅ 629M ⋅ 42M ⋅ 42.8 | 12th ⋅ 360H ⋅ 821M ⋅ 59M ⋅ 57.3 |
| seat | Y5 | Y10 | Y15 | Y20 |
|---|---|---|---|---|
| claude-fable-5 | 2nd ⋅ 180H ⋅ 326M ⋅ 105M ⋅ 47.4 | 5th ⋅ 240H ⋅ 394M ⋅ 196M ⋅ 56.4 | 5th ⋅ 440H ⋅ 568M ⋅ 250M ⋅ 71.3 | 3rd ⋅ 620H ⋅ 507M ⋅ 298M ⋅ 76.3 |
| muse-spark-1.1 | 8th ⋅ 0H ⋅ 260M ⋅ 84M ⋅ 18.5 | 7th ⋅ 40H ⋅ 197M ⋅ 117M ⋅ 24.0 | 2nd ⋅ 160H ⋅ 272M ⋅ 222M ⋅ 46.2 | 1st ⋅ 420H ⋅ 306M ⋅ 246M ⋅ 62.5 |
| deepseek-v4-pro | 9th ⋅ 0H ⋅ 279M ⋅ 91M ⋅ 20.0 | 10th ⋅ 140H ⋅ 342M ⋅ 91M ⋅ 44.5 | 4th ⋅ 200H ⋅ 397M ⋅ 122M ⋅ 52.2 | 5th ⋅ 200H ⋅ 363M ⋅ 160M ⋅ 52.1 |
| glm-5.2 | 4th ⋅ 40H ⋅ 275M ⋅ 107M ⋅ 30.3 | 1st ⋅ 160H ⋅ 335M ⋅ 78M ⋅ 46.1 | 8th ⋅ 200H ⋅ 346M ⋅ 71M ⋅ 49.3 | 6th ⋅ 200H ⋅ 356M ⋅ 125M ⋅ 51.4 |
| grok-4.5 | 7th ⋅ 20H ⋅ 197M ⋅ 134M ⋅ 21.1 | 6th ⋅ 260H ⋅ 145M ⋅ 145M ⋅ 40.0 | 7th ⋅ 320H ⋅ 108M ⋅ 185M ⋅ 36.4 | 8th ⋅ 520H ⋅ 139M ⋅ 199M ⋅ 50.9 |
| gpt-5.6-sol | 6th ⋅ 100H ⋅ 258M ⋅ 115M ⋅ 36.9 | 4th ⋅ 120H ⋅ 239M ⋅ 86M ⋅ 36.8 | 1st ⋅ 240H ⋅ 281M ⋅ 229M ⋅ 52.3 | 2nd ⋅ 280H ⋅ 182M ⋅ 294M ⋅ 47.9 |
| minimax-m3 | 3rd ⋅ 120H ⋅ 272M ⋅ 93M ⋅ 38.2 | 11th ⋅ 220H ⋅ 246M ⋅ 61M ⋅ 43.2 | 13th ⋅ 220H ⋅ 236M ⋅ 64M ⋅ 42.5 | 13th ⋅ 220H ⋅ 260M ⋅ 76M ⋅ 44.8 |
| claude-sonnet-5 | 11th ⋅ 100H ⋅ 246M ⋅ 82M ⋅ 35.9 | 13th ⋅ 100H ⋅ 268M ⋅ 64M ⋅ 36.4 | 9th ⋅ 100H ⋅ 251M ⋅ 77M ⋅ 36.0 | 10th ⋅ 160H ⋅ 226M ⋅ 94M ⋅ 40.8 |
| kimi-k2.6 | 13th ⋅ 0H ⋅ 238M ⋅ 76M ⋅ 18.1 | 8th ⋅ 20H ⋅ 305M ⋅ 82M ⋅ 26.9 | 3rd ⋅ 60H ⋅ 351M ⋅ 106M ⋅ 36.9 | 7th ⋅ 80H ⋅ 304M ⋅ 175M ⋅ 39.7 |
| gpt-5.6-terra | 5th ⋅ 40H ⋅ 266M ⋅ 108M ⋅ 28.8 | 2nd ⋅ 60H ⋅ 113M ⋅ 108M ⋅ 13.0 | 6th ⋅ 60H ⋅ 352M ⋅ 108M ⋅ 36.2 | 9th ⋅ 60H ⋅ 390M ⋅ 138M ⋅ 38.4 |
| qwen3.7-max | 14th ⋅ 20H ⋅ 279M ⋅ 87M ⋅ 19.2 | 12th ⋅ 20H ⋅ 254M ⋅ 88M ⋅ 13.6 | 12th ⋅ 180H ⋅ 451M ⋅ 200M ⋅ 34.2 | 11th ⋅ 180H ⋅ 451M ⋅ 113M ⋅ 32.8 |
| claude-opus-4.8 | 10th ⋅ 20H ⋅ 274M ⋅ 90M ⋅ 18.4 | 3rd ⋅ 40H ⋅ 263M ⋅ 99M ⋅ 21.2 | 10th ⋅ 40H ⋅ 412M ⋅ 132M ⋅ 27.8 | 14th ⋅ 40H ⋅ 472M ⋅ 95M ⋅ 28.4 |
| gemini-3-flash | 1st ⋅ 140H ⋅ −6M ⋅ 110M ⋅ 12.7 | 9th ⋅ 160H ⋅ −77M ⋅ 107M ⋅ 10.7 | 11th ⋅ 160H ⋅ 92M ⋅ 111M ⋅ 11.7 | 4th ⋅ 200H ⋅ 107M ⋅ 182M ⋅ 15.8 |
| claude-haiku-4.5 | 16th ⋅ 0H ⋅ 209M ⋅ 64M ⋅ 11.0 | 14th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 4.3 | 14th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 0.8 | 16th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 0.8 |
| gemini-3.5-flash | 12th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 2.1 | 15th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.2 | 15th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.2 | 12th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.2 |
| tool | what it does |
|---|---|
| Query (12), charged to the 30-query budget: | |
| get_club_overview | the agent’s club at a glance: division, cash, board, facilities, tactics |
| get_squad | the senior squad with ability bands, form, contracts, wages |
| get_player | detailed view of any player by id |
| get_league_table | league standings |
| get_fixtures | the club’s season fixtures and results |
| get_match_report | match report by match id |
| get_transfer_market | players available to buy: listed, free agents, expiring contracts |
| get_finances | cash, revenue, wage bill, active warnings |
| get_youth_academy | academy prospects with potential stars |
| get_history | the run’s archive: honors, seasons, transfers |
| get_inbox | pending offers and open items |
| get_draft_pool | draft phase only: the shared template pool with scouted ability bands, potential stars, prices, and preset contracts, plus budget status |
| Action (11); negotiation moves charged to the 10-move budget: | |
| set_lineup | set the preferred starting XI (auto-completed if players are unavailable) |
| set_tactics | set formation and playing style from fixed enums |
| make_transfer_offer | bid for another club’s player; the negotiation resolves within the stop |
| respond_to_offer | accept, reject, counter, or accept a counter on a pending offer |
| offer_contract | renew an own player or sign a free agent at a wage and term |
| list_player | put a player on or off the transfer list |
| release_player | terminate a contract (severance: half of one year’s wage) |
| promote_youth | promote an academy prospect to the senior squad |
| invest | start a facility upgrade: academy, training, or stadium |
| set_standing_order | add or clear a standing order (max 10; one threshold plus a fixed action) |
| submit_draft | draft phase only: submit the complete pick list; resubmitting replaces it, and auto-fill completes a short list from the cheapest tier |
| Notebook (2), free: | |
| append_note | append to the private notebook (persists across stops) |
| rewrite_notes | replace the entire notebook |
| Control (1): | |
| advance | end this decision stop and simulate to the next one |
| stop (day of the 360-day year) | what the world wants decided |
|---|---|
| Scheduled, 13 per season: | |
| Preseason board (2) | the board sets the season target; plan the year |
| Summer window plan (5) | the transfer window opens; build the squad |
| Summer deadline (59) | last actions before the window closes |
| Lineup lock (63) | commit lineup and tactics before round 1 |
| Monthly report (90, 120, 150, 240) | finances and form check-in |
| Winter window open (181) | mid-course squad correction |
| Midseason review (190) | the board measures progress against target |
| Winter deadline (209) | last winter-window actions |
| Season settlement (300) | the season closes: honors, revenue, board verdict |
| Youth intake (305) | the academy class arrives; promote, hold, or release |
| Event-driven, unsolicited: | |
| Draft (once, at run start) | assemble the 25-player squad from the shared pool |
| Offer batch | incoming bids for the agent’s players, batched at a minimum 15-day spacing |
| Contract expiry | a senior contract is running down; renew or lose the player for free |
| Major injury | a first-team player goes down; cover or reshuffle |
| Board warning | confidence is deteriorating; the board expects a response |
| Insolvency warning | cash runway is shrinking (30- and 60-day warnings) |
| Administration | insolvency executed: points penalty and forced sales |
| Revival | the seat restarts under the capped-revival mechanism (§4.2) |
| layer | who writes it | contents | delivery |
|---|---|---|---|
| Packet auto-context | engine | dashboard digest, inbox (all events since last stop), agent’s last 10 engine-mutating actions, pending offers | pushed free into every stop packet |
| Archive | engine | honors / season / transfer history of the whole run | get_history query tool, on demand |
| Notebook | the agent | intentions, reasoning, plans (“why I bought him”, “sell Y in winter”) | injected into every packet; written via append_note / rewrite_notes |
| Audit log | runner | every successful action | never fed back: replay/anti-cheat/resume only |
| quantity | median | range |
|---|---|---|
| decision stops | 374 | 341–390 |
| of which scheduled | ∼260 (13/season) | rest event-triggered (Table 11) |
| matches simulated in the world | 4,800 | (30 rounds × 8 matches × 20 seasons) |
| model turns (API calls) | ∼1,760 | 1,184–6,969 |
| state-changing actions | — | 849–2,198 |
| tokens (manifest accounting) | 61M | 24M–274M |
| notebook writes | 354 | 320–425 |
| notebook size at year 20 | — | 0.4k–209k characters |
| simulation parameters (frozen) | 313 | — |
왜 중요한가
짧은 문제 하나를 잘 푸는 것과 수년에 걸쳐 누적되는 결정을 잘 관리하는 것은 전혀 다른 능력이라는 점을 실제로 측정 가능하게 만들었다는 데 의미가 있다. 앞으로 에이전트를 회사 운영, 장기 투자, 자원 관리처럼 실제 세계의 '경영' 업무에 투입하려는 사람들에게, 모델 크기나 가격표가 아니라 어떤 행동 습관을 봐야 하는지에 대한 구체적인 기준을 제시한다.
이 논문의 용어
- 에이전트(agent) · 스스로 도구를 골라 실행하며 목표를 수행하는 AI 시스템
- 솔로 트랙 / 아레나 · 솔로 트랙은 모델 혼자 고정된 상대와 겨루는 모드, 아레나는 여러 모델이 하나의 세계를 공유하며 직접 경쟁하는 모드
- 오라클(oracle) · 숨겨진 정보를 모두 볼 수 있는 특권을 가진 비교용 스크립트, 최고 성능의 기준선 역할
- 노트북(notebook) · 매 순간 대화가 초기화되는 환경에서 에이전트가 스스로 남기는 유일한 기록, 기억 유지 전략 자체가 평가 대상
- Sfinal(최종 점수) · 우승 등 성과 점수, 순자산 증가분, 스쿼드 가치를 합쳐 20년 전체를 누적 평가하는 단일 점수
본문에 싣지 못한 그림
- Figure 1: How a run works. One run spans 20 in-game years and roughly 340 to 400 decision stops, ending in a single composite score. At each stop, the clock freezes and the agent takes as many tool-call turns as it wants (read → act → advance). The agent’s only cross-stop memory is its self-written notebook (§3.2); the environment side holds the club (net worth, squad with permanently hidden true ability, a board that can fire the manager) and the world (15 rival clubs, a counter-adaptive market, delayed-payoff investments).
- Figure 2: Per-season solo trajectories for all 15 models in four channels. Lines are the mean over three seeds, shading spans the best and worst seed, and the five models with the highest mean are colored and named in the legend; the grey lines are the remaining ten models.
- Figure 3: Credit assignment. Left: discretionary spend bucketed by payoff horizon, with run totals at right. Right: mean season-end idle-cash ratio vs. final score, rs=−0.50, negative on every seed.
- Figure 4: Proactive control. Left: distribution of contract months remaining when each renewal negotiation episode opens (rows sorted by final score; shaded band = last-minute, ≤6 months). Right: per-model median lead vs. final score, rs=+0.45, positive on every seed.
- Figure 5: Capability matrix over the 15 solo models, six axes. Each column rank-normalizes one behavioral metric averaged over the three seeds (1 = best of 15; construction in Appendix 15), and rows are sorted by mean final score. The winner is a generalist, with no axis below 0.79 and mean 0.94, rather than the leader of any single one; the mid-table pairs a genuine strength with a decisive gap; and the bottom of the board is weak on most axes at once.
- Figure 6: Arena trajectories, all 16 seats. The six highest-scoring seats are colored and named in the legend; the grey lines are the remaining ten seats. Left: composite Sfinal at each season end. Right: league-table position.
- Figure 7: Final score per model, mean with standard deviation over the three seeds, individual seeds shown as points. Dashed and dotted lines mark the oracle and heuristic means. Stability separates models that a single-seed board cannot tell apart, from kimi-k2.6 at 0.15 to claude-haiku-4.5 at 22.73.
- Figure 8: Transfer offers per season, solo track (rows sorted by final score; totals at right). The winner’s 9 offers are all in the first two seasons; the field-wide offer-to-completion conversion is 2–5%.
- Figure 9: Transfer offers per season in the Arena (approximate year mapping; daggered models settled early). The winner’s offer count rises from 9 (solo) to 91: the same model under a different, correct judgment of market liquidity.
- Figure 10: Season-by-season cumulative API cost vs. composite score, solo track. Steep curves convert dollars into score throughout; flat long curves do not. The full 20-year runs span $18 to $191.
- Figure 11: The failed metric: share of stops carrying an engine warning, by type (left) and against final score (right), means over the three seeds. The per-seed coefficient runs +0.45 / +0.07 / −0.78 and averages to nothing, so warning exposure is not reported as a capability. Zero warnings conflates genuine anticipation with do-nothing conservatism.
- Figure 12: Both panels are means over the three seeds. Left: youth harvest rate (share of promotions later fielded in ≥5 lineups); the correlation that looked informative on seed 1 does not survive three (−0.52 / −0.18 / +0.04). Right: endgame shift in long-horizon actions per season, years 2–16 (blue) vs. 17–20 (red); endgame-aware models reduce late investment, one ramps up.
- Figure 13: Memory-curation regimes on the solo track. Left: season-to-season TF–IDF cosine similarity of each model’s reconstructed season-end notebook. Right: notebook size over the run. The right panel decodes the left: the same “high consistency” is a 200k-character append-only archive for one model and a 3–6k curated document for the winner.
- Figure 14: Compute allocation. Mean tokens per run against mean solo score, both over the three seeds, on a log axis. A null result (rs=−0.19, p=0.50) across a sevenfold spend range.
- Figure 15: The asset channels behind Figure 6: per-season net worth and squad value for all 16 Arena seats, the six highest-scoring seats colored as in Figure 6 and the grey lines the remaining ten (numeric snapshots in Table 9). The winner’s mid-table seasons are asset accumulation (the net-worth curve keeps climbing while the standings do not), and the conservative cash-holders’ net worth grows on cash the composite discounts.
논문 원문 초록 (영문)
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget a
arXiv에서 원문 보기최신 논문
- FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI SystemsAI가 '정답'을 맞혀도, 정작 증거는 못 찾는 경우가 대부분이었다 - 재무 이상거래 진단 AI의 숨은 약점
- FACET: Preserving Source Intent and Executable State in Terminal Task SynthesisAI 에이전트에게 '터미널 문제집'을 정합성 있게 만들어주는 파이프라인, FACET
- Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models음성·소리를 알아듣는 AI, 지시문 학습 없이 '연결 다리'만 훈련해도 충분하다
- Adversarial Review: Structured Disagreement for Grounded Agentic Code ReviewAI 코딩 에이전트, 에이전트 늘리기보다 '검토자 vs 비판자' 한 쌍이 더 똑똑하게 일한다
- Looped Language Models Improve Compositional Tool Calling생각을 여러 번 되짚는 AI가 여러 개의 도구를 순서대로 엮어 쓰는 일도 더 잘한다
- Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement쇼핑몰 검색에서 이탈한 고객을 AI 에이전트가 다시 카톡(왓츠앱)으로 불러온 이야기
- GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence Networks바이러스 유전자 서열을 코돈끼리 서로 옆에 등장하는 관계망(그래프)으로 바꿔 변이를 구분하는 법
- Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval이미지-긴문장 검색 AI가 '이미 다 맞혔다'고 착각해서 정작 헷갈리는 문제를 못 배우던 버릇을 고쳤다