Mathematics in the age of AI
If AI can eventually solve research-level math proofs, what should the math community actually protect?
This essay, based on a lecture at the 2026 International Congress of Mathematicians, sidesteps the debate over whether AI can already do research-level math and instead assumes it eventually will. It then asks what mathematicians actually value and optimize for, using problem solving as a case study. The author argues that simply producing a correct answer is not enough — understanding, exposition, and community acceptance matter just as much.
What they did
- Instead of arguing about current AI math capability, the author adopts a 'Working Hypothesis' that AI will eventually handle a reasonable share of research-level math tasks, and asks what mathematical values would need re-examining as a result.
- Using problem solving as a case study, the goal is refined step by step: from 'solve as many problems as possible' to adding correctness verification, clear communication, community digestion/acceptance, and finally incorporation into a field's definitive, canonical theory.
- The essay warns that an AI-generated proof can be formally verified yet understood by nobody, or so smoothly polished that the genuinely difficult steps become indistinguishable from routine ones, erasing cues that normally guide readers.
- As concrete evidence, the second batch of the First Proof project tested four AI systems on ten genuinely new research-level problems never posted online; seven of the ten received at least one passing, expert-refereed grade, at compute costs of roughly tens to hundreds of dollars per problem.
- The piece quotes several recommendations from the June 2026 Leiden Declaration on Artificial Intelligence and Mathematics — endorsed by the International Mathematical Union — on disclosing AI tool use, supporting reviewers, affirming human authorship, and proper attribution.
![Figure 3. A page from a 1991 paper of Bourgain [3], annotated by my much younger (and very frustrated) self. But by fighting my way through these texts, I came to understand Bourgain’s way of thinking, and in time I actively sought out his papers to read. See also [19], [17].](https://media.metallab.ai/papers/2608.16753/f0.jpg)
Why it matters
As AI starts generating proofs faster than the community can verify, write up, referee, or canonicalize them, existing institutions built for scarcity of proofs may break down under abundance. The essay is a reminder — relevant to anyone thinking about knowledge production, not just mathematicians — that a fast, correct answer is not the same as understood, trusted knowledge.
Terms in this paper
- Goodhart's law · When a measure becomes a target, it stops being a reliable measure of what it was meant to track
- autoformalization · automatically converting human-written math proofs into machine-checkable proof assistant languages like Lean
- canonicalization · the slow process of restating a proven result in its natural, general form and absorbing it into standard textbooks and toolkits
- First Proof project · an independent evaluation testing AI systems on brand-new, never-published research-level math problems
- Leiden Declaration · a set of 23 recommendations on AI and mathematics, published June 2026 and endorsed by the International Mathematical Union
Original abstract (English)
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
Read on arXivLatest papers
- Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal AbstractionsA small language for scheduling problems lets weaker AI models write feasible schedules instead of broken code
- SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured DecompositionMaking AI judges show their work when picking the better of two answers
- Position: AI Leaderboards Are Underserving the Global South: A Case Study from IndiaIndia and other Global South regions already have solid AI benchmarks, but no trusted referee to rank results fairly
- Temporal Multi-Signal Fusion for Token-Level Hallucination DetectionCatching AI's made-up facts works better when you read them as a stretch of text, not word by word
- Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector ApplicationOff-the-shelf open-source AI fails 3 out of 4 times on a high-risk public-sector document task
- When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress ClassificationA stress-detecting wearable AI averaged 93% accuracy but totally failed on one person, so researchers built a pre-check that flags risky readings before the AI even makes a guess
- Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-EngagementAI agents win back window-shopping customers by chasing them down on WhatsApp
- Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical LessonYou don't need to generate explanation text on every request - pre-make it and just pick one
Latest from METAL LAB
- Open-source Ornith-1.5 claims scores on par with Claude Opus
- Modular, Now Owned by Qualcomm, Fully Open-Sources Mojo Language
- Anthropic Launches Free "Academy" Site for Learning Claude
- Liquid AI's 300M Draft Model Speeds Up Decoding by Up to 3.18x
- ChatGPT sites now handled by Codex instead of Git and CI, adds team invites
Figures: jonbaer et al., arXiv:2608.16753, arxiv-nonexclusive