One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Epoch AI unveils undisclosed game puzzle benchmark — Opus 5 scores 59%

Following up on its Chess Puzzles benchmark, Epoch AI measures reasoning ability with a puzzle set from an undisclosed game

이미지: METAL LAB 생성

Summary

  • Epoch AI has released a new 'game puzzles' benchmark structured similarly to its existing Chess Puzzles benchmark.
  • The identity of the game the puzzles come from has not been disclosed, with the goal of measuring reasoning tasks that models are unlikely to have encountered through post-training.
  • As of the announcement, the top score reported is 59% by Opus 5.
발표 주체
Epoch AI
발표 시점
2026년 8월 6일(UTC)
벤치마크 명칭
game puzzles
설계 참조
Epoch AI의 기존 Chess Puzzles 벤치마크
사용 게임
비공개(undisclosed)
현재 1위
Opus 5
1위 점수
59%
측정 목적
사후학습이 이뤄지지 않았을 가능성이 큰 추론 중심 과제 평가

After chess puzzles, an unnamed game

Epoch AI has released a new 'game puzzles' benchmark. The evaluation follows a format similar to the organization's previously operated Chess Puzzles benchmark, but draws problems from a different game rather than chess — one whose identity has not been disclosed.

Epoch AI tied its reason for withholding the game's identity to the possibility of post-training exposure. The concern is that puzzle formats from well-known games may have already been repeatedly covered during model training and tuning processes, making it difficult to measure pure reasoning ability. Epoch AI stated that the benchmark tests reasoning-focused tasks that models are "unlikely to have been post-trained on."

Image of Epoch AI's released game puzzles benchmark results
Game puzzles benchmark announcement image · Source: Epoch AI

Current top score: Opus 5 at 59%

As of the announcement, the record-holding model is Opus 5, with a score of 59%. Epoch AI did not provide individual scores for other models or a full leaderboard alongside this announcement, so gaps between top-tier models and comparisons against human baselines remain unconfirmed.

ItemChess Puzzles (existing)game puzzles (new)
Source gameChessUndisclosed
Design approachPuzzle-based evaluationSimilar structure applied
Disclosed scoresNot included in this announcementOpus 5: 59%
Modelgame puzzles score
Opus 5bar:59 59%

Why an "uncontaminated" evaluation is needed

Reasoning benchmarks have long faced criticism that once problem types become public, they can be absorbed into subsequent model training. The approach of keeping the game's identity itself undisclosed appears designed to delay this kind of exposure. However, how long Epoch AI plans to keep this undisclosed, along with specifics such as the number of problems and scoring methodology, cannot be determined from this announcement alone.

Ongoing debate over AI model evaluation methods and benchmark reliability can be followed in METAL LAB's benchmark-related articles.

What remains to be confirmed

This report is based on a single public post from Epoch AI. The full benchmark leaderboard, the list of comparison models, and the evaluation protocol will need to be confirmed through additional disclosures.