
이미지: The Decoder
Summary
- Fields Medalist Terence Tao warned in an essay for the 2026 International Congress of Mathematicians (ICM) that AI could bring turmoil to mathematics on par with the foundational crisis of the early 1900s
- He argued that this time, what's being tested isn't mathematical truth itself but the unspoken value system that determines what counts as an achievement and who gets credit for it
- He cited an experiment from the First-Proof project in which 7 of 10 unpublished research problems received a passing verdict from at least one AI system
- 발표
- 테렌스 타오, 2026 국제수학자대회(ICM) 제출 에세이
- 비교 대상
- 1900~1930년 수학 기초 위기(러셀의 역설, 괴델 불완전성 정리)
- 실험 근거
- 퍼스트-프루프 프로젝트 2차 라운드, 미공개 연구문제 10개 중 7개가 AI 4개 시스템 중 최소 1곳에서 통과 판정
- 실험 비용
- 문제당 수십~수백 달러
- 검증 지연 사례
- 에르되시 문제 데이터베이스에 검증 안 된 AI 생성 제출물 수십 건 누적
- 지침 문서
- 2026년 6월 공개, 국제수학연맹 지지 '라이덴 선언'
- 타오의 기준
- 저자가 전문가 수준으로 명확히 설명하지 못하는 결과는 게재 보류해야 함
- 타오의 AI 활용
- 문헌 검색, 도표 제작, 텍스트 완성, 슬라이드의 논문 변환에 한정
The nature of the crisis has changed
In an essay submitted to the 2026 International Congress of Mathematicians (ICM), Fields Medalist Terence Tao warned that AI could trigger the biggest crisis in mathematics since the early 20th century. Rather than fixating on the question of "what can AI do," he wrote, the mathematical community needs to confront questions it has long deferred: what counts as a research contribution, who deserves credit for it, and what it even means to say something has been "understood."
Tao compared the situation to the foundational crisis that shook mathematics between 1900 and 1930. Back then, Russell's paradox and Gödel's incompleteness theorems forced mathematicians to lay bare assumptions they had tacitly taken for granted, ultimately producing a rigorous framework that has held up for a century. What's being tested now, Tao noted, isn't the coherence of truth but a set of practices that were never explicitly codified in the first place — practices that define what counts as a contribution, what gets rewarded, and whether a machine can be said to have "done the work."
Evidence that AI is already solving problems
Tao summarized his working hypothesis this way: "AI tools will soon be able to carry out a substantial share of research-level mathematical tasks, at a reasonable success rate, quality, level of supervision, and cost." As evidence, he pointed to the First-Proof project. In its second round of testing, four AI systems were given ten previously unpublished research problems under controlled conditions — and seven of the ten received a passing grade from at least one system, producing solutions that were either essentially flawless or required only minor fixes, at a cost of tens to a few hundred dollars per problem.
Tao pointed out that mathematics's various goals — solving problems, building theory, forming community, training the next generation — have traditionally been tightly interwoven, and that AI risks severing these connections. To explain the risk, he invoked Goodhart's Law: once a metric becomes the target, it stops being a good measure. He argued that generative AI is especially vulnerable to this trap because it tends to chase outputs that look good rather than outputs that actually deliver results. The investment structure of the AI industry compounds the problem, since it rewards precisely the citable, benchmark-measurable outputs that mathematicians have long used as proxies for deeper goals.
Proofs are piling up faster than verification can keep pace
If Tao's hypothesis holds, he warned, mathematics could shift from an era of proof scarcity to an era of proof glut — with AI-generated proofs accumulating faster than anyone could check, read, or digest them all. He pointed out that the Erdős Problems database already contains dozens of AI-generated submissions that no expert has stepped forward to verify.
AI-polished proofs carry a further problem. Human-written proofs naturally retain a certain friction at their hardest points — a carefully constructed lemma, traces of changed notation, a paragraph that has obviously been rewritten several times. An overly polished AI-generated proof erases this noise along with the signal, producing writing that's "easy to read but hard to learn from." As Tao put it, "the 'mistakes' in human-written exposition can actually be helpful to the reader."
The Leiden Declaration and Tao's own standard
As a concrete guideline, Tao pointed to the Leiden Declaration, released in June 2026 and endorsed by the International Mathematical Union. His own standard is unambiguous: "unless the authors can convincingly demonstrate that they can give a clear, expert-level presentation of their results, with correct and proper attribution, that result should not be published." Even a formally verified proof, in his view, should be considered incomplete if no human can properly explain it.
He also stressed that training young mathematicians requires special care — that mathematicians need to preserve the "irreducible human aspect" of their work and strictly limit their use of AI tools, since producing homework with correct answers is not the same as training a mathematician. Tao noted that he himself limits his own AI use to literature searches, generating diagrams, completing text, and converting slides into paper format.
Tao's essay strikes a different note from the recent debate between fellow Fields Medalist Timothy Gowers and mathematician Peter Sarnak, who argued that LLMs possess real mathematical ability but hit their limits when genuinely novel ideas are required. Where Gowers and Sarnak asked "what can AI do," Tao is asking "what does the field of mathematics itself actually want."
Editor's take
What makes this essay unsettling isn't that AI's ability is worrying — it's that AI's ability is worrying precisely because it's middling. If AI failed completely, we could ignore it; if it succeeded perfectly, we'd just need to verify it. The real problem lies in the zone the First-Proof experiment revealed: plausibly solving seven out of ten problems. What happens at that point isn't a decline in quality but an explosion in volume, as evidenced by the pile of unverified submissions accumulating in the Erdős database. When humans do the verifying but machines do the producing, the bottleneck is guaranteed to burst at verification.
Anyone with long experience in peer review will find this structure familiar. Reviewer shortages have long been a chronic ailment of academia — but when submission volume suddenly triples or quadruples, that chronic condition turns acute. That's why it matters that Tao's proposed solution isn't a benchmark score but a return to something like an oral exam: can the author explain their own result in front of experts? It amounts to routing the verification of machine-made output back through a distinctly human capacity — the ability to explain.
The same problem is coming soon to Korea's mathematics and science community. It won't be long before dissertation defenses, conference presentations, and grant reviews all need some procedure for asking "how much of this result did AI produce." The preparation that's possible now is straightforward: set rules in advance that confine AI tool use to auxiliary tasks like literature search and drafting, while requiring that authors themselves be able to orally defend the core proofs or arguments. In the coming months, it's likely that conferences and journals will begin rolling out AI-use disclosure policies modeled on guidelines like the Leiden Declaration.




Comments