Leading mathematicians say large language models have become genuinely useful at mathematics but still cannot make the leaps that produce new theory.

Timothy Gowers and Peter Sarnak, two of the most decorated names in the field, credit current models with real skill at combining established methods and exploring many lines of attack. In a blog post on August 12, 2026, Gowers wrote that today's models are good at recombining known techniques and testing many search paths, but lack the intuition to pick the few productive routes inside a vast space of possibilities.
Sarnak, writing in the July 2026 issue of the AMS Notices, reaches a similar verdict. AI can derive results from existing theory, he argues, yet fails to develop the abstractions that underpin major proofs when it starts from an elementary question.
DeepMind researcher Tom Zahavy framed the limit in a paper titled "LLMs Can't Jump." He locates the bottleneck in what he calls "manipulative abduction," the ability to invent new foundational assumptions that have no precedent in language. World models, he suggests, could be a path past the ceiling.
The assessments feed a wider argument over whether LLMs are becoming broadly more capable or simply getting better at benchmarks and familiar problem types. For now, the mathematicians agree the machine is a powerful calculator that has not yet learned to surprise them.
Sources
- The Decoder: "Top mathematicians say LLMs are strong calculators but poor creative thinkers" (https://the-decoder.com/top-mathematicians-say-llms-are-strong-calculators-but-poor-creative-thinkers/)
- Timothy Gowers: "What sort of maths are LLMs good at?" (https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-are-llms-good-at/)
- Peter Sarnak, AMS Notices, July 2026 (https://www.ams.org/journals/notices/202607/noti3373/noti3373.html)
- Tom Zahavy, "LLMs Can't Jump" (https://the-decoder.com/language-models-cant-spark-scientific-revolutions-but-world-models-might/)



