
이미지: METAL LAB 생성
Summary
- Epoch AI has released a new 'game puzzles' benchmark structured similarly to its existing Chess Puzzles benchmark.
- The identity of the game the puzzles come from has not been disclosed, with the goal of measuring reasoning tasks that models are unlikely to have encountered through post-training.
- As of the announcement, the top score reported is 59% by Opus 5.
- 발표 주체
- Epoch AI
- 발표 시점
- 2026년 8월 6일(UTC)
- 벤치마크 명칭
- game puzzles
- 설계 참조
- Epoch AI의 기존 Chess Puzzles 벤치마크
- 사용 게임
- 비공개(undisclosed)
- 현재 1위
- Opus 5
- 1위 점수
- 59%
- 측정 목적
- 사후학습이 이뤄지지 않았을 가능성이 큰 추론 중심 과제 평가
After chess puzzles, an unnamed game
Epoch AI has released a new 'game puzzles' benchmark. The evaluation follows a format similar to the organization's previously operated Chess Puzzles benchmark, but draws problems from a different game rather than chess — one whose identity has not been disclosed.
Epoch AI tied its reason for withholding the game's identity to the possibility of post-training exposure. The concern is that puzzle formats from well-known games may have already been repeatedly covered during model training and tuning processes, making it difficult to measure pure reasoning ability. Epoch AI stated that the benchmark tests reasoning-focused tasks that models are "unlikely to have been post-trained on."

Current top score: Opus 5 at 59%
As of the announcement, the record-holding model is Opus 5, with a score of 59%. Epoch AI did not provide individual scores for other models or a full leaderboard alongside this announcement, so gaps between top-tier models and comparisons against human baselines remain unconfirmed.
| Item | Chess Puzzles (existing) | game puzzles (new) |
|---|---|---|
| Source game | Chess | Undisclosed |
| Design approach | Puzzle-based evaluation | Similar structure applied |
| Disclosed scores | Not included in this announcement | Opus 5: 59% |
| Model | game puzzles score |
|---|---|
| Opus 5 | bar:59 59% |
Why an "uncontaminated" evaluation is needed
Reasoning benchmarks have long faced criticism that once problem types become public, they can be absorbed into subsequent model training. The approach of keeping the game's identity itself undisclosed appears designed to delay this kind of exposure. However, how long Epoch AI plans to keep this undisclosed, along with specifics such as the number of problems and scoring methodology, cannot be determined from this announcement alone.
Ongoing debate over AI model evaluation methods and benchmark reliability can be followed in METAL LAB's benchmark-related articles.
What remains to be confirmed
This report is based on a single public post from Epoch AI. The full benchmark leaderboard, the list of comparison models, and the evaluation protocol will need to be confirmed through additional disclosures.



