One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

OpenAI's Math Model Solves 10 Open Problems, Sparks Citation Dispute

Hundreds of pages of supporting documents released, but a key attribution line was quietly edited — mathematicians are divided

노란 배경 위에 놓인 회색 주판 3D 렌더 이미지

이미지: The Verge AI

Summary

  • OpenAI announced that an unnamed internal model solved 10 unsolved math problems, including quantum game theory and high-dimensional sphere packing
  • Initially describing the problems as having seen "no progress in a decade," OpenAI later quietly revised the claim to acknowledge it built on the prior work of two researchers, triggering a citation controversy
  • It also emerged that the same model family still makes errors on simple tasks like calculating days of the week or basic arithmetic
발표 주체
오픈AI
발표 형식
블로그 글 '10 Advances in Mathematics and Theoretical Computer Science', 수백 쪽 분량 증빙 문서 첨부
해결 주장 분야
양자 게임이론, 3차원을 넘는 고차원 구 채우기 등 10개 분야
사용 모델
이름 미공개 내부 모델(추정 아스트라), 오픈AI가 확인은 거부
이전 사례
지난 5월 오픈AI 내부 모델이 80년 된 '단위 거리 추측' 반증
논란
'10년간 진전 없음' 서술을 이후 두 연구자 기존 연구 인용으로 조용히 수정
취재
로버트 하트, 더버지 런던 주재 AI 기자

An announcement backed by hundreds of pages of evidence

A blog post OpenAI published in August shook the math world. Titled "10 Advances in Mathematics and Theoretical Computer Science," it claimed to have solved problems across ten different areas of mathematics and theoretical computer science. To back up the claim, OpenAI posted hundreds of pages of supporting documentation alongside it. Robert Hart, The Verge's London-based AI reporter, covered the announcement and interviewed several mathematicians about it — some of that reporting also appeared in the article A Fields Medalist Grapples With His Identity as AI Conquers Mathematics.

OpenAI did not disclose the model's name. Appearing on the podcast Decoder, Hart said he believed the model was likely Astral, but that OpenAI declined to confirm this when he asked. Back in May, an unnamed internal OpenAI model had also disproved the "unit distance conjecture," which had gone unsolved for 80 years, and Hart suspected this was the same model family.

From quantum game theory to sphere packing

The list of 10 problems in the announcement included quantum game theory and sphere-packing problems in dimensions higher than three. None of these are lightweight topics. In his interview, Hart quoted one researcher as saying that "solving even one of these would have guaranteed an academic career." Unlike past cases where AI models' math achievements were criticized for focusing on narrow problems that the academic community didn't care much about, these ten were problems that multiple mathematicians had genuinely worked on for a long time — which is part of why reactions were mixed.

One quietly changed sentence

The core of the controversy was attribution. In its initial announcement, OpenAI stated there had been no meaningful progress on these 10 problems in the past decade. But the actual paper clearly stated that one of the problems had been solved building on work by two existing researchers. This portion was later quietly revised without explanation. Some researchers Hart contacted expressed frustration that OpenAI had not adequately credited prior work that had, in fact, contributed substantially. However, Hart noted that one of the named researchers took an ambiguous stance on the matter. Hart characterized the issue not as plagiarism but as a case of a sloppy press release overstating its achievements.

Still can't count days of the week

Separate from these advanced mathematical achievements, the same model family still struggles with basic tasks. Reports continue to surface of errors on simple tasks like calculating days of the week or tracking time. Hart pointed out that mathematics isn't a single unified field. Just as biology spans everything from animal observation to cellular biochemistry, mathematics spans a wide spectrum from simple calculation to high-level abstract theory. AI models are strong at combining existing methods in new ways and connecting disparate domains, but they still show weaknesses when it comes to handling precise values, like counting.

Editor's View

What stands out in this controversy is the fact that OpenAI specifically emphasized the number "10." Announcing each problem individually would take longer to verify and generate less buzz. Bundling ten together lets OpenAI make a big impression before the press and academia can fully scrutinize each claim. Placed alongside Anthropic's August 10 announcement of partial progress on the Riemann hypothesis using an unreleased research version of Claude, a clearer picture emerges: frontier model developers are competitively using mathematics as a verifiable stage to showcase their capabilities.

This echoes points previously raised by Fields Medalist Timothy Gowers and mathematician Peter Sarnak — that LLMs are skilled at combining existing methods across a vast search space, but lack the intuition to judge which paths are actually productive. Many of these ten problems were extensions or continuations of existing research, and the fact that this was initially downplayed reveals the same limitation. Generating entirely new abstractions is a different task from rapidly assembling existing pieces into an answer.

From a practical standpoint, it's clear what academia and research institutions should do now. When reviewing math achievement announcements from AI labs, the habit should be to check the attached supporting documents and citation lists for prior work before looking at the headline results. Judging solely by the numbers in a press release makes it easy to miss quietly revised claims like this one. For math departments and research funding bodies, this actually strengthens the case for continued investment in discovering new problems and mentoring talent — because the ability to quickly solve existing problems and the ability to pose new questions remain fundamentally different skills.

In the coming months, frontier labs including OpenAI and Anthropic will likely continue rolling out math achievement announcements. For future announcements, expect academia to push back with calls for pre-verification or co-author confirmation processes, so that attribution issues like this one aren't quietly revised again.

Comments