
Image: METAL
Summary
- Fields Medalist Timothy Gowers and mathematician Peter Sarnak have pointed to the limits of LLMs' mathematical creativity
- Gowers diagnosed that LLMs lack the intuition to pick productive paths through a vast search space
- Google DeepMind researcher Tom Zahavy reached the same conclusion in his paper "LLMs Can't Jump"
Computation, yes. Ideas, no.
Proving a new theorem is less about computation and more about conceiving "a first step no one has tried before." That's the point at the heart of two separate assessments: one from Fields Medalist Timothy Gowers on his own blog, the other from number theory expert Peter Sarnak in an interview with the American Mathematical Society's (AMS) Notices. Both acknowledge that large language models (LLMs) show considerable ability in mathematics, but argue there's a fundamental limit to their capacity to generate genuinely new ideas.
In a blog post published on the 12th, Gowers wrote that current models are strong at combining already-known methods and pursuing multiple search paths simultaneously. The problem comes after that. He pointed out that models lack the intuition to identify, within a vast search space, which paths are actually productive — the "this direction is right" sense that human mathematicians develop only after decades of training.
Sarnak's explanation in the AMS Notices interview is somewhat more concrete. He said AI can derive results from already-established theories, but cannot start from a very basic question and independently invent the new abstract concepts that underpin major proofs. The weight of this point becomes clearer when you consider how the great theorems of number theory first emerged — not as the product of skillfully applying existing tools, but as the creation of concepts that did not previously exist at all.
A DeepMind researcher reaches the same conclusion
A similar diagnosis has come from within the industry itself. Google DeepMind researcher Tom Zahavy gave this bottleneck a name in his paper "LLMs Can't Jump." He frames the core issue with a concept he calls "manipulative abduction" — the ability to independently invent new foundational assumptions where no linguistic precedent exists at all. This is distinct from rearranging existing sentences or extending patterns; it is, literally, creating something that did not exist before.
Zahavy suggested world models — AI models that learn the structure of the actual environment, such as visual and physical information, rather than language — as a candidate for overcoming this limitation. The logic is that if language-bound learning is the problem, a model that learns structure outside of language directly could be an alternative. Whether this proposal actually solves the problem of mathematical creativity, however, remains unverified.
Benchmarks keep rising, so why doesn't creativity?
This diagnosis connects to a question that has been repeating across the industry recently: LLMs keep setting new records on various test scores, but the capability actually felt in practical work and research settings hasn't grown as much. In the PerceptionBench results released by Moonshot AI on the 15th, none of 16 frontier models exceeded 60% accuracy on pure visual perception, with the top model reaching only 59.7%. This showed that models celebrated for their reasoning ability falter even at the basic stage of accurately reading an image.
Gowers and Sarnak's points run in the same direction: there is a gap between benchmark scores and actual creative ability, and that gap doesn't close on its own just because models get bigger. That said, both mathematicians acknowledge LLMs' computational and combinatorial abilities, so reading their remarks as "AI is useless" would be an exaggeration. What they pointed to is a limitation in one specific domain — creating fundamentally new mathematical concepts.
Editor's view
What makes these remarks interesting is that neither speaker is an AI skeptic. Gowers has long leaned toward the view that AI can assist mathematical research, and Sarnak, too, doesn't deny AI's usefulness as a computational tool. The fact that these two, speaking separately around the same time, both raised the "wall of creativity" reads as a signal that even optimists are starting to see this wall.
Anyone who has actually put LLMs to work will find this familiar. Models are remarkably fast and accurate at applying known theorems or filling in specific steps of a proof. But reframing a problem — seeing it through an entirely different lens from the start — rarely emerges. Something similar shows up in code review or draft reports: models quickly combine existing patterns to produce plausible results, but proposals that overturn the problem definition itself are rare. This is likely not because models are lazy, but because the data needed to learn that kind of reframing is itself scarce to begin with.
The lesson for domestic research and education is clear. Using LLMs as an "assistant that fills in specific proof steps" remains valid and efficient today. But expecting them to "suggest new research directions" is still overreach. Rather than adopting AI brainstorming outright in graduate seminars or R&D planning stages, limiting its use to quickly combining and verifying known methods reduces the chance of failure.
In the coming months, this debate is likely to intensify alongside announcements of world-model research. If models that learn structure outside of language, as Zahavy proposed, actually show something close to "manipulative abduction," the two mathematicians' diagnosis could turn out to be a provisional conclusion that doesn't last long. Conversely, if such attempts stall, the frame of "LLMs are calculators" is likely to settle in as the industry's default for the time being.





Comments