
이미지: METAL LAB 생성
Summary
- Anthropic announced on August 10 that it had an unreleased research version of Claude attempt to solve the Riemann Hypothesis.
- While it did not prove the hypothesis, the company claimed it raised the existing lower bound on the proportion of zeros of the Riemann zeta function that satisfy the hypothesis.
- The excerpted post did not include specific figures for the improvement before and after, so the scale of the achievement cannot be judged from the original text alone.
- 발표 주체
- 앤트로픽(Anthropic) 공식 X 계정
- 발표 시점
- 2026년 8월 10일
- 사용 모델
- 외부 미공개 연구용 Claude 버전
- 결과
- 리만 가설 자체는 해결하지 못함
- 부분 성과
- 제타 함수 영점 중 가설을 만족하는 비율의 하한 상향
- 수치
- 개선 전후 값은 공개된 발췌문에 포함되지 않음
A 160-year-old problem posed to an unreleased Claude
Anthropic had an as-yet-unreleased research version of Claude attempt to solve the Riemann Hypothesis. To cut to the result: it didn't solve it. Instead, the company said it "made progress on a related problem" — specifically, raising the lower bound on the proportion of zeros of the Riemann zeta function that satisfy the hypothesis. However, the publicly excerpted post cuts off at "in," so it's not clear exactly what figure was raised or by how much.
Why the Riemann Hypothesis is so hard
The Riemann Hypothesis is one of mathematics' most famous unsolved problems, proposed by Bernhard Riemann in 1859 and still unresolved today. It concerns how prime numbers are spaced along the number line, and is one of the seven Millennium Prize Problems for which the Clay Mathematics Institute has offered a $1 million reward.
Here's the core idea: there's a function called the zeta function, and it has points where it equals zero. Riemann conjectured that all of these points, without exception, lie on a single line in the complex plane — the so-called "critical line." Computers have calculated trillions of zeros, and not a single counterexample has ever turned up. Yet the hypothesis remains unproven. Showing that "all" of the infinitely many zeros lie on that line is a task that no amount of case-checking can ever finish.
A workaround: raising the "percentage" instead of proving "all"
So mathematicians took a workaround. Instead of proving that all zeros lie on the line, they calculate what minimum percentage of zeros do, and gradually push that proportion higher. In 1974, Norman Levinson raised it to at least one-third; in 1989, Brian Conrey pushed it to two-fifths. Since then, improvements measured in decimal points have each been enough to produce a paper. It's a field where even the halfway mark hasn't been crossed in half a century. What Anthropic describes as "raising the lower bound" belongs squarely to this lineage of work.
This is the important part. When AI has been applied to mathematics before, it has mostly been to verify human-made proofs or solve benchmark problems. This case, by contrast, is a claim of incremental progress on an actual research frontier that professional mathematicians have been working on for decades. Of course, such a result only counts as an achievement once it's published as a paper and reviewed by other mathematicians. There's a long distance between a single tweet and academic recognition.
AMD invests up to $7 trillion won in Anthropic, to supply 2GW of chips
So what does this change
Anthropic is the company behind Claude, and has recently been growing its business around coding agents and enterprise models. It's not hard to guess why such a company would bring up an "unreleased research version" to talk about an unsolved problem. It's a signal that the stage on which AI models demonstrate their capability is shifting from test scores to unsolved problems. Google DeepMind's results on math olympiad problems and sorting algorithms sit on the same trend line.
For readers, this isn't an announcement that adds something immediately usable — the model in question hasn't been released. But there's something else worth watching. If frontier labs' future achievement announcements start being framed not as "benchmark score X" but as "how far we pushed on which unsolved problem," the very yardstick for comparing AI capability changes. Results from real, verifiable problems are harder to inflate than benchmark scores, which could make them a more honest indicator for readers in the end.



