FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
Handing an AI a football club to run for 20 years reveals that winning comes from management habits, not raw model power
FM-Bench is a benchmark where a language model agent manages a simulated football club for 20 in-game years, using 26 tools across roughly 340 to 400 decision stops. It builds in hidden player abilities, delayed-payoff investments, a transfer market that adapts to the agent's behavior, and a board judging results and finances together, to test whether models can sustain good decisions over a long horizon. Across 15 frontier models run three times each in both a solo mode and a shared 'Arena' mode, claude-fable-5 topped both tracks, yet model scale, price, and vendor did not predict the overall order.
What they did
- Built a 20-year simulated football-management environment to test whether language model agents can keep making good decisions when actions have effects that compound over a long time.
- Instantiated four hard aspects of real management: player ability that is never fully knowable, investments that only pay off years later, a transfer market that raises its hidden prices in response to rejected offers, and a board that judges both results and financial discipline.
- Tested 15 frontier models (closed models from Anthropic, OpenAI, Google, xAI, Meta, plus five open-weight models) in a solo track against a fixed scripted world and in an Arena where all 15 compete inside one shared 20-year world.
- Broke the single final score down into six behavioral capabilities, such as cutting long-payoff investment near the end of the run, avoiding idle cash, and renewing player contracts well before deadlines, to see which habits actually predict success.
- claude-fable-5 won both tracks (solo mean 90.94 versus 95.54 for a scripted policy that can see hidden information), but in the Arena the league title still rotated among ten different models across seasons.
| demand | mechanism | targeted ability | calibration lever |
|---|---|---|---|
| Hidden information | scout bands with a permanent per-scout bias; hidden player traits (injury proneness, development rate, aging onset); hidden asks in negotiation | valuation calibration under noise | bandwidth, bias sigma |
| Cumulative consequences | Upside: youth development and facility investment pay off over years. Downside: insolvency ends in administration (points penalty, forced sales), confidence collapse ends in firing; both compound yet remain recoverable. Honors accrue into the final score season by season | long-range credit assignment; trend recognition, loss-cutting | growth and aging curves; spiral thresholds, warning cadence; honors weights |
| Counter-adaptive market | rejected bids raise the hidden ask; repeat-pair markups (bargains included); anti-inversion counters; negotiation cooldowns; per-seed mispricing | strategy adaptation | markup and decay rates |
| Multi-objective pressure | the board judges results and financial discipline jointly; season targets scale with squad strength | multi-constraint balancing | target cushion, board patience |
| seat | Sfinal | tokens |
|---|---|---|
| oracle (privileged) | 95.54±4.68 | — |
| claude-fable-5 | 90.94±5.20 | 24M |
| kimi-k2.6 | 88.49±0.15 | 87M |
| gpt-5.6-terra | 86.66±1.20 | 28M |
| gpt-5.6-sol | 86.40±2.53 | 62M |
| muse-spark-1.1 | 83.19±12.03 | 73M |
| glm-5.2 | 83.17±0.92 | 58M |
| grok-4.5 | 81.82±9.87 | 191M |
| qwen3.7-max | 80.66±10.38 | 47M |
| deepseek-v4-pro | 79.15±2.67 | 194M |
| gemini-3-flash | 79.06±13.45 | 51M |
| claude-sonnet-5 | 75.75±5.03 | 39M |
| claude-opus-4.8 | 75.02±2.46 | 31M |
| gemini-3.5-flash | 74.59±11.02 | 154M |
| minimax-m3 | 68.37±12.33 | 74M |
| claude-haiku-4.5 | 36.90±22.73 | 86M |
| heuristic | 17.05±12.34 | — |
| idle | −0.90±1.86 | — |
| random | −17.21±2.45 | — |
| player | Sfinal | settle | deaths | stops | time |
|---|---|---|---|---|---|
| H1 | 74.64 | completed | 0 | 394 | 10 h |
| H2 | 59.95 | completed | 0 | 398 | 4 h |
| H3 | 5.39 | fired, t=10.8 | 4 | 202 | 4 h |
| H4 | 2.47 | fired, t=10.7 | 4 | 198 | 5 h |
| H5 | 0.41 | fired, t=7.8 | 4 | 165 | 1.5 h |
| H6 | −3.78 | fired, t=12.3 | 4 | 237 | 2 h |
| # | seat | Sfinal | deaths | settle | tokens |
|---|---|---|---|---|---|
| 1 | claude-fable-5 | 76.26 | 0 | completed | ∼34M |
| 2 | muse-spark-1.1 | 62.47 | 0 | completed | ∼80M |
| 3 | deepseek-v4-pro | 52.09 | 0 | completed | ∼111M |
| 4 | glm-5.2 | 51.44 | 0 | completed | ∼86M |
| 5 | grok-4.5 | 50.89 | 0 | completed | ∼138M |
| 6 | gpt-5.6-sol | 47.94 | 0 | completed | ∼78M |
| 7 | minimax-m3 | 44.78 | 0 | completed | ∼65M |
| 8 | claude-sonnet-5 | 40.75 | 0 | completed | ∼42M |
| 9 | kimi-k2.6 | 39.72 | 0 | completed | ∼119M |
| 10 | gpt-5.6-terra | 38.44 | 0 | completed | ∼19M |
| 11 | qwen3.7-max | 32.77 | 2 | completed | ∼54M |
| 12 | claude-opus-4.8 | 28.44 | 1 | completed | ∼43M |
| 13 | gemini-3-flash | 15.76 | 2 | completed | ∼63M |
| 14 | claude-haiku-4.5 | 0.76 | 4 | fired, t=10.7 | ∼30M |
| 15 | gemini-3.5-flash | 0.21 | 4 | fired, t=5.8 | ∼40M |
| 16 | heuristic (anchor) | 0.13 | 4 | fired, t=2.6 | — |
| # | seat | provider | role |
|---|---|---|---|
| 1 | claude-fable-5 (4) | Anthropic | new flagship |
| 2 | claude-opus-4.8 (5) | Anthropic | previous flagship |
| 3 | claude-haiku-4.5 (3) | Anthropic | small tier |
| 4 | gpt-5.6-sol (32) | OpenAI | new flagship |
| 5 | gpt-5.6-terra (32) | OpenAI | balanced tier |
| 6 | claude-sonnet-5 (6) | Anthropic | workhorse mid (4-point family curve) |
| 7 | gemini-3.5-flash (16) | latest GA flagship | |
| 8 | gemini-3-flash (15) | previous-gen fast | |
| 9 | deepseek-v4-pro (40) | Together | open flagship |
| 10 | grok-4.5 (38) | xAI | frontier |
| 11 | kimi-k2.6 (30) | Together | open agentic specialist |
| 12 | qwen3.7-max (1) | Together | open flagship |
| 13 | glm-5.2 (42) | Together | open top tier |
| 14 | muse-spark-1.1 (27) | Meta Model API | Meta agentic model |
| 15 | minimax-m3 (29) | Together | open flagship |
| 16 | heuristic | — (scripted) | disciplined-script anchor |
| lever | easy | medium | hard |
|---|---|---|---|
| scouting noise, multiplier on the belief sigma | 1.3 | 1.0 | 0.35 |
| lineup rotates for fatigue | no | yes | yes |
| market days acted on | every 2nd | every | every |
| gap required before buying an upgrade | 1.5 | 1.0 | 0.25 |
| wage budget, multiplier on the board target | 0.92 | 1.0 | 1.10 |
| seat | score | sd | profile |
|---|---|---|---|
| claude-fable-5 | 90.9 | 5.2 | generalist, no axis below 0.79; best bidder in the field at 9 offers per signing and the sharpest endgame reduction |
| kimi-k2.6 | 88.5 | 0.1 | the steadiest model in the campaign, scoring within a third of a point on three different worlds |
| gpt-5.6-terra | 86.7 | 1.2 | minimalist: lowest token spend in the field (28M) after the winner and an early endgame reduction, but the shortest renewal leads among the top models |
| gpt-5.6-sol | 86.4 | 2.5 | archivist memory (similarity 0.91, append-only) and the weakest bidding of the top group (45 offers per signing) |
| muse-spark-1.1 | 83.2 | 12.0 | strong on two worlds and 20 points weaker on the third; one of two models that raise long-horizon spending at the end |
| glm-5.2 | 83.2 | 0.9 | stable across seeds but loose with cash (135% idle ratio) |
| grok-4.5 | 81.8 | 9.9 | compute as substitute: long renewal leads and heavy spend (191M tokens) with wide seed-to-seed swings |
| qwen3.7-max | 80.7 | 10.4 | longest renewal leads after the winner, undone by a 103% idle ratio and churning memory (similarity 0.23) |
| deepseek-v4-pro | 79.2 | 2.7 | near-static notebook (0.81) and the heaviest spend in the field (194M tokens) |
| gemini-3-flash | 79.1 | 13.5 | best cash discipline in the field (46%) but the least willing to cut endgame spending; collapses on the hardest seed |
| claude-sonnet-5 | 75.7 | 5.0 | churning memory (0.20, the lowest) and no endgame reduction |
| claude-opus-4.8 | 75.0 | 2.5 | shortest renewal leads in the field (10 months) with a 134% idle ratio |
| gemini-3.5-flash | 74.6 | 11.0 | worst price discovery (73 offers per signing) and the largest endgame ramp-up |
| minimax-m3 | 68.4 | 12.3 | prefers veteran signings, short leads, and swings 23 points across seeds |
| claude-haiku-4.5 | 36.9 | 22.7 | weak on every axis at once: 196% idle cash, 11-month leads, a rising endgame rate, and the widest spread in the campaign |
| seat | Y5 | Y10 | Y15 | Y20 |
|---|---|---|---|---|
| claude-fable-5 | 4th ⋅ 140H ⋅ 293M ⋅ 130M ⋅ 33.3 | 1st ⋅ 480H ⋅ 470M ⋅ 218M ⋅ 50.3 | 1st ⋅ 900H ⋅ 1048M ⋅ 443M ⋅ 65.9 | 1st ⋅ 1400H ⋅ 1742M ⋅ 634M ⋅ 94.6 |
| muse-spark-1.1 | 3rd ⋅ 120H ⋅ 302M ⋅ 158M ⋅ 31.1 | 1st ⋅ 460H ⋅ 468M ⋅ 196M ⋅ 48.7 | 1st ⋅ 960H ⋅ 1037M ⋅ 457M ⋅ 66.5 | 1st ⋅ 1460H ⋅ 1298M ⋅ 315M ⋅ 91.2 |
| grok-4.5 | 4th ⋅ 60H ⋅ 289M ⋅ 169M ⋅ 23.9 | 1st ⋅ 460H ⋅ 318M ⋅ 107M ⋅ 46.6 | 1st ⋅ 880H ⋅ 618M ⋅ 225M ⋅ 58.1 | 1st ⋅ 1380H ⋅ 1106M ⋅ 634M ⋅ 89.7 |
| kimi-k2.6 | 3rd ⋅ 120H ⋅ 344M ⋅ 116M ⋅ 31.9 | 1st ⋅ 440H ⋅ 430M ⋅ 99M ⋅ 46.7 | 1st ⋅ 940H ⋅ 668M ⋅ 276M ⋅ 61.4 | 1st ⋅ 1440H ⋅ 733M ⋅ 400M ⋅ 88.5 |
| gemini-3-flash | 3rd ⋅ 160H ⋅ 247M ⋅ 165M ⋅ 32.8 | 1st ⋅ 500H ⋅ 339M ⋅ 111M ⋅ 42.1 | 1st ⋅ 1000H ⋅ 675M ⋅ 400M ⋅ 61.5 | 1st ⋅ 1420H ⋅ 812M ⋅ 618M ⋅ 86.6 |
| qwen3.7-max | 3rd ⋅ 120H ⋅ 404M ⋅ 127M ⋅ 34.0 | 1st ⋅ 360H ⋅ 582M ⋅ 140M ⋅ 48.0 | 1st ⋅ 780H ⋅ 917M ⋅ 256M ⋅ 62.3 | 1st ⋅ 1200H ⋅ 1305M ⋅ 248M ⋅ 86.1 |
| gpt-5.6-terra | 3rd ⋅ 140H ⋅ 359M ⋅ 120M ⋅ 34.0 | 1st ⋅ 440H ⋅ 565M ⋅ 102M ⋅ 50.1 | 1st ⋅ 780H ⋅ 808M ⋅ 247M ⋅ 59.8 | 1st ⋅ 1280H ⋅ 987M ⋅ 186M ⋅ 85.3 |
| gpt-5.6-sol | 4th ⋅ 140H ⋅ 343M ⋅ 135M ⋅ 33.6 | 1st ⋅ 460H ⋅ 615M ⋅ 176M ⋅ 51.5 | 1st ⋅ 960H ⋅ 762M ⋅ 217M ⋅ 62.6 | 2nd ⋅ 1220H ⋅ 1109M ⋅ 274M ⋅ 84.5 |
| glm-5.2 | 3rd ⋅ 120H ⋅ 294M ⋅ 121M ⋅ 31.8 | 1st ⋅ 420H ⋅ 513M ⋅ 110M ⋅ 48.1 | 3rd ⋅ 760H ⋅ 712M ⋅ 58M ⋅ 56.6 | 1st ⋅ 1080H ⋅ 1331M ⋅ 266M ⋅ 82.8 |
| minimax-m3 | 10th ⋅ 20H ⋅ 324M ⋅ 154M ⋅ 20.9 | 1st ⋅ 240H ⋅ 492M ⋅ 100M ⋅ 39.3 | 1st ⋅ 540H ⋅ 827M ⋅ 155M ⋅ 51.8 | 1st ⋅ 1040H ⋅ 1245M ⋅ 327M ⋅ 82.5 |
| claude-sonnet-5 | 5th ⋅ 20H ⋅ 247M ⋅ 91M ⋅ 18.6 | 2nd ⋅ 180H ⋅ 363M ⋅ 97M ⋅ 36.0 | 1st ⋅ 480H ⋅ 647M ⋅ 94M ⋅ 50.2 | 1st ⋅ 980H ⋅ 1020M ⋅ 146M ⋅ 80.4 |
| deepseek-v4-pro | 5th ⋅ 140H ⋅ 385M ⋅ 124M ⋅ 35.8 | 1st ⋅ 480H ⋅ 638M ⋅ 111M ⋅ 52.9 | 1st ⋅ 820H ⋅ 732M ⋅ 110M ⋅ 58.7 | 3rd ⋅ 1080H ⋅ 996M ⋅ 163M ⋅ 79.0 |
| claude-opus-4.8 | 3rd ⋅ 140H ⋅ 308M ⋅ 100M ⋅ 32.7 | 1st ⋅ 280H ⋅ 467M ⋅ 78M ⋅ 40.5 | 2nd ⋅ 600H ⋅ 788M ⋅ 160M ⋅ 56.4 | 3rd ⋅ 760H ⋅ 1313M ⋅ 216M ⋅ 76.9 |
| gemini-3.5-flash | 2nd ⋅ 140H ⋅ 260M ⋅ 186M ⋅ 29.9 | 1st ⋅ 480H ⋅ 226M ⋅ 140M ⋅ 41.3 | 1st ⋅ 980H ⋅ 479M ⋅ 156M ⋅ 57.2 | 5th ⋅ 1040H ⋅ 877M ⋅ 170M ⋅ 75.2 |
| claude-haiku-4.5 | 12th ⋅ 100H ⋅ 293M ⋅ 132M ⋅ 28.4 | 7th ⋅ 200H ⋅ 395M ⋅ 39M ⋅ 33.8 | 3rd ⋅ 340H ⋅ 629M ⋅ 42M ⋅ 42.8 | 12th ⋅ 360H ⋅ 821M ⋅ 59M ⋅ 57.3 |
| seat | Y5 | Y10 | Y15 | Y20 |
|---|---|---|---|---|
| claude-fable-5 | 2nd ⋅ 180H ⋅ 326M ⋅ 105M ⋅ 47.4 | 5th ⋅ 240H ⋅ 394M ⋅ 196M ⋅ 56.4 | 5th ⋅ 440H ⋅ 568M ⋅ 250M ⋅ 71.3 | 3rd ⋅ 620H ⋅ 507M ⋅ 298M ⋅ 76.3 |
| muse-spark-1.1 | 8th ⋅ 0H ⋅ 260M ⋅ 84M ⋅ 18.5 | 7th ⋅ 40H ⋅ 197M ⋅ 117M ⋅ 24.0 | 2nd ⋅ 160H ⋅ 272M ⋅ 222M ⋅ 46.2 | 1st ⋅ 420H ⋅ 306M ⋅ 246M ⋅ 62.5 |
| deepseek-v4-pro | 9th ⋅ 0H ⋅ 279M ⋅ 91M ⋅ 20.0 | 10th ⋅ 140H ⋅ 342M ⋅ 91M ⋅ 44.5 | 4th ⋅ 200H ⋅ 397M ⋅ 122M ⋅ 52.2 | 5th ⋅ 200H ⋅ 363M ⋅ 160M ⋅ 52.1 |
| glm-5.2 | 4th ⋅ 40H ⋅ 275M ⋅ 107M ⋅ 30.3 | 1st ⋅ 160H ⋅ 335M ⋅ 78M ⋅ 46.1 | 8th ⋅ 200H ⋅ 346M ⋅ 71M ⋅ 49.3 | 6th ⋅ 200H ⋅ 356M ⋅ 125M ⋅ 51.4 |
| grok-4.5 | 7th ⋅ 20H ⋅ 197M ⋅ 134M ⋅ 21.1 | 6th ⋅ 260H ⋅ 145M ⋅ 145M ⋅ 40.0 | 7th ⋅ 320H ⋅ 108M ⋅ 185M ⋅ 36.4 | 8th ⋅ 520H ⋅ 139M ⋅ 199M ⋅ 50.9 |
| gpt-5.6-sol | 6th ⋅ 100H ⋅ 258M ⋅ 115M ⋅ 36.9 | 4th ⋅ 120H ⋅ 239M ⋅ 86M ⋅ 36.8 | 1st ⋅ 240H ⋅ 281M ⋅ 229M ⋅ 52.3 | 2nd ⋅ 280H ⋅ 182M ⋅ 294M ⋅ 47.9 |
| minimax-m3 | 3rd ⋅ 120H ⋅ 272M ⋅ 93M ⋅ 38.2 | 11th ⋅ 220H ⋅ 246M ⋅ 61M ⋅ 43.2 | 13th ⋅ 220H ⋅ 236M ⋅ 64M ⋅ 42.5 | 13th ⋅ 220H ⋅ 260M ⋅ 76M ⋅ 44.8 |
| claude-sonnet-5 | 11th ⋅ 100H ⋅ 246M ⋅ 82M ⋅ 35.9 | 13th ⋅ 100H ⋅ 268M ⋅ 64M ⋅ 36.4 | 9th ⋅ 100H ⋅ 251M ⋅ 77M ⋅ 36.0 | 10th ⋅ 160H ⋅ 226M ⋅ 94M ⋅ 40.8 |
| kimi-k2.6 | 13th ⋅ 0H ⋅ 238M ⋅ 76M ⋅ 18.1 | 8th ⋅ 20H ⋅ 305M ⋅ 82M ⋅ 26.9 | 3rd ⋅ 60H ⋅ 351M ⋅ 106M ⋅ 36.9 | 7th ⋅ 80H ⋅ 304M ⋅ 175M ⋅ 39.7 |
| gpt-5.6-terra | 5th ⋅ 40H ⋅ 266M ⋅ 108M ⋅ 28.8 | 2nd ⋅ 60H ⋅ 113M ⋅ 108M ⋅ 13.0 | 6th ⋅ 60H ⋅ 352M ⋅ 108M ⋅ 36.2 | 9th ⋅ 60H ⋅ 390M ⋅ 138M ⋅ 38.4 |
| qwen3.7-max | 14th ⋅ 20H ⋅ 279M ⋅ 87M ⋅ 19.2 | 12th ⋅ 20H ⋅ 254M ⋅ 88M ⋅ 13.6 | 12th ⋅ 180H ⋅ 451M ⋅ 200M ⋅ 34.2 | 11th ⋅ 180H ⋅ 451M ⋅ 113M ⋅ 32.8 |
| claude-opus-4.8 | 10th ⋅ 20H ⋅ 274M ⋅ 90M ⋅ 18.4 | 3rd ⋅ 40H ⋅ 263M ⋅ 99M ⋅ 21.2 | 10th ⋅ 40H ⋅ 412M ⋅ 132M ⋅ 27.8 | 14th ⋅ 40H ⋅ 472M ⋅ 95M ⋅ 28.4 |
| gemini-3-flash | 1st ⋅ 140H ⋅ −6M ⋅ 110M ⋅ 12.7 | 9th ⋅ 160H ⋅ −77M ⋅ 107M ⋅ 10.7 | 11th ⋅ 160H ⋅ 92M ⋅ 111M ⋅ 11.7 | 4th ⋅ 200H ⋅ 107M ⋅ 182M ⋅ 15.8 |
| claude-haiku-4.5 | 16th ⋅ 0H ⋅ 209M ⋅ 64M ⋅ 11.0 | 14th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 4.3 | 14th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 0.8 | 16th ⋅ 0H ⋅ 181M ⋅ 19M ⋅ 0.8 |
| gemini-3.5-flash | 12th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 2.1 | 15th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.2 | 15th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.2 | 12th ⋅ 0H ⋅ 192M ⋅ 105M ⋅ 0.2 |
| tool | what it does |
|---|---|
| Query (12), charged to the 30-query budget: | |
| get_club_overview | the agent’s club at a glance: division, cash, board, facilities, tactics |
| get_squad | the senior squad with ability bands, form, contracts, wages |
| get_player | detailed view of any player by id |
| get_league_table | league standings |
| get_fixtures | the club’s season fixtures and results |
| get_match_report | match report by match id |
| get_transfer_market | players available to buy: listed, free agents, expiring contracts |
| get_finances | cash, revenue, wage bill, active warnings |
| get_youth_academy | academy prospects with potential stars |
| get_history | the run’s archive: honors, seasons, transfers |
| get_inbox | pending offers and open items |
| get_draft_pool | draft phase only: the shared template pool with scouted ability bands, potential stars, prices, and preset contracts, plus budget status |
| Action (11); negotiation moves charged to the 10-move budget: | |
| set_lineup | set the preferred starting XI (auto-completed if players are unavailable) |
| set_tactics | set formation and playing style from fixed enums |
| make_transfer_offer | bid for another club’s player; the negotiation resolves within the stop |
| respond_to_offer | accept, reject, counter, or accept a counter on a pending offer |
| offer_contract | renew an own player or sign a free agent at a wage and term |
| list_player | put a player on or off the transfer list |
| release_player | terminate a contract (severance: half of one year’s wage) |
| promote_youth | promote an academy prospect to the senior squad |
| invest | start a facility upgrade: academy, training, or stadium |
| set_standing_order | add or clear a standing order (max 10; one threshold plus a fixed action) |
| submit_draft | draft phase only: submit the complete pick list; resubmitting replaces it, and auto-fill completes a short list from the cheapest tier |
| Notebook (2), free: | |
| append_note | append to the private notebook (persists across stops) |
| rewrite_notes | replace the entire notebook |
| Control (1): | |
| advance | end this decision stop and simulate to the next one |
| stop (day of the 360-day year) | what the world wants decided |
|---|---|
| Scheduled, 13 per season: | |
| Preseason board (2) | the board sets the season target; plan the year |
| Summer window plan (5) | the transfer window opens; build the squad |
| Summer deadline (59) | last actions before the window closes |
| Lineup lock (63) | commit lineup and tactics before round 1 |
| Monthly report (90, 120, 150, 240) | finances and form check-in |
| Winter window open (181) | mid-course squad correction |
| Midseason review (190) | the board measures progress against target |
| Winter deadline (209) | last winter-window actions |
| Season settlement (300) | the season closes: honors, revenue, board verdict |
| Youth intake (305) | the academy class arrives; promote, hold, or release |
| Event-driven, unsolicited: | |
| Draft (once, at run start) | assemble the 25-player squad from the shared pool |
| Offer batch | incoming bids for the agent’s players, batched at a minimum 15-day spacing |
| Contract expiry | a senior contract is running down; renew or lose the player for free |
| Major injury | a first-team player goes down; cover or reshuffle |
| Board warning | confidence is deteriorating; the board expects a response |
| Insolvency warning | cash runway is shrinking (30- and 60-day warnings) |
| Administration | insolvency executed: points penalty and forced sales |
| Revival | the seat restarts under the capped-revival mechanism (§4.2) |
| layer | who writes it | contents | delivery |
|---|---|---|---|
| Packet auto-context | engine | dashboard digest, inbox (all events since last stop), agent’s last 10 engine-mutating actions, pending offers | pushed free into every stop packet |
| Archive | engine | honors / season / transfer history of the whole run | get_history query tool, on demand |
| Notebook | the agent | intentions, reasoning, plans (“why I bought him”, “sell Y in winter”) | injected into every packet; written via append_note / rewrite_notes |
| Audit log | runner | every successful action | never fed back: replay/anti-cheat/resume only |
| quantity | median | range |
|---|---|---|
| decision stops | 374 | 341–390 |
| of which scheduled | ∼260 (13/season) | rest event-triggered (Table 11) |
| matches simulated in the world | 4,800 | (30 rounds × 8 matches × 20 seasons) |
| model turns (API calls) | ∼1,760 | 1,184–6,969 |
| state-changing actions | — | 849–2,198 |
| tokens (manifest accounting) | 61M | 24M–274M |
| notebook writes | 354 | 320–425 |
| notebook size at year 20 | — | 0.4k–209k characters |
| simulation parameters (frozen) | 313 | — |
Why it matters
It turns 'can an AI agent make one good decision' into something measurable over years of compounding consequences, which is closer to how real organizations actually operate. For anyone considering deploying agents for long-running tasks like running a company, managing investments, or allocating resources, it shows that behavioral habits matter more than model size or price tag.
Terms in this paper
- agent · an AI system that autonomously chooses and uses tools to accomplish a task
- solo track / Arena · solo track runs one model against a fixed scripted world; Arena places multiple models in one shared world where they directly compete
- oracle · a scripted reference policy with privileged access to hidden information, used as a soft upper-bound benchmark
- notebook · the agent's self-written memory, the only information carried between decision points since each stop starts a fresh conversation
- Sfinal (final score) · the single cumulative score combining honors won, net-worth growth, and squad value over the full 20-year run
Figures we cannot republish
- Figure 1: How a run works. One run spans 20 in-game years and roughly 340 to 400 decision stops, ending in a single composite score. At each stop, the clock freezes and the agent takes as many tool-call turns as it wants (read → act → advance). The agent’s only cross-stop memory is its self-written notebook (§3.2); the environment side holds the club (net worth, squad with permanently hidden true ability, a board that can fire the manager) and the world (15 rival clubs, a counter-adaptive market, delayed-payoff investments).
- Figure 2: Per-season solo trajectories for all 15 models in four channels. Lines are the mean over three seeds, shading spans the best and worst seed, and the five models with the highest mean are colored and named in the legend; the grey lines are the remaining ten models.
- Figure 3: Credit assignment. Left: discretionary spend bucketed by payoff horizon, with run totals at right. Right: mean season-end idle-cash ratio vs. final score, rs=−0.50, negative on every seed.
- Figure 4: Proactive control. Left: distribution of contract months remaining when each renewal negotiation episode opens (rows sorted by final score; shaded band = last-minute, ≤6 months). Right: per-model median lead vs. final score, rs=+0.45, positive on every seed.
- Figure 5: Capability matrix over the 15 solo models, six axes. Each column rank-normalizes one behavioral metric averaged over the three seeds (1 = best of 15; construction in Appendix 15), and rows are sorted by mean final score. The winner is a generalist, with no axis below 0.79 and mean 0.94, rather than the leader of any single one; the mid-table pairs a genuine strength with a decisive gap; and the bottom of the board is weak on most axes at once.
- Figure 6: Arena trajectories, all 16 seats. The six highest-scoring seats are colored and named in the legend; the grey lines are the remaining ten seats. Left: composite Sfinal at each season end. Right: league-table position.
- Figure 7: Final score per model, mean with standard deviation over the three seeds, individual seeds shown as points. Dashed and dotted lines mark the oracle and heuristic means. Stability separates models that a single-seed board cannot tell apart, from kimi-k2.6 at 0.15 to claude-haiku-4.5 at 22.73.
- Figure 8: Transfer offers per season, solo track (rows sorted by final score; totals at right). The winner’s 9 offers are all in the first two seasons; the field-wide offer-to-completion conversion is 2–5%.
- Figure 9: Transfer offers per season in the Arena (approximate year mapping; daggered models settled early). The winner’s offer count rises from 9 (solo) to 91: the same model under a different, correct judgment of market liquidity.
- Figure 10: Season-by-season cumulative API cost vs. composite score, solo track. Steep curves convert dollars into score throughout; flat long curves do not. The full 20-year runs span $18 to $191.
- Figure 11: The failed metric: share of stops carrying an engine warning, by type (left) and against final score (right), means over the three seeds. The per-seed coefficient runs +0.45 / +0.07 / −0.78 and averages to nothing, so warning exposure is not reported as a capability. Zero warnings conflates genuine anticipation with do-nothing conservatism.
- Figure 12: Both panels are means over the three seeds. Left: youth harvest rate (share of promotions later fielded in ≥5 lineups); the correlation that looked informative on seed 1 does not survive three (−0.52 / −0.18 / +0.04). Right: endgame shift in long-horizon actions per season, years 2–16 (blue) vs. 17–20 (red); endgame-aware models reduce late investment, one ramps up.
- Figure 13: Memory-curation regimes on the solo track. Left: season-to-season TF–IDF cosine similarity of each model’s reconstructed season-end notebook. Right: notebook size over the run. The right panel decodes the left: the same “high consistency” is a 200k-character append-only archive for one model and a 3–6k curated document for the winner.
- Figure 14: Compute allocation. Mean tokens per run against mean solo score, both over the three seeds, on a log axis. A null result (rs=−0.19, p=0.50) across a sevenfold spend range.
- Figure 15: The asset channels behind Figure 6: per-season net worth and squad value for all 16 Arena seats, the six highest-scoring seats colored as in Figure 6 and the grey lines the remaining ten (numeric snapshots in Table 9). The winner’s mid-table seasons are asset accumulation (the net-worth curve keeps climbing while the standings do not), and the conservative cash-holders’ net worth grows on cash the composite discounts.
Original abstract (English)
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget a
Read on arXivLatest papers
- FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI SystemsAI often names the right cause of a financial mismatch without ever finding the proof for it
- Looped Language Models Improve Compositional Tool CallingAI models that rethink their own answers multiple times get better at chaining tools together
- FACET: Preserving Source Intent and Executable State in Terminal Task SynthesisFACET builds internally consistent terminal-task 'exam sets' to train command-line AI agents
- Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-EngagementAI agents win back window-shopping customers by chasing them down on WhatsApp
- Adversarial Review: Structured Disagreement for Grounded Agentic Code ReviewFor AI code review, one reviewer plus one critic beats piling on more agents
- GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence NetworksTurning viral gene sequences into codon relationship maps to tell coronavirus variants apart
- Alignment Is All You Need: Instruction-Free Training for General Audio-Language ModelsA frozen language model plus one lightweight connector is enough to build a capable audio-understanding AI
- Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption RetrievalAn image-to-long-caption search AI kept thinking it had already solved the problem, so it never learned to tell near-identical captions apart
Latest from METAL LAB
- NVIDIA's 300 Verified Skills Lift Correctness by 41 Points
- Wave your hand at a webcam, hear a theremin: browser instrument released
- Meta AI launches desktop app for Mac, can read an entire app window
- Factory Commits $100M to Partner Network, Pushes to Scale Software Factories
- SpaceX approached Cognition for acquisition four days after closing Cursor deal