Recent advances in large language models and proof assistants have enabled AI systems to generate, formalize, and verify complex mathematical proofs, raising new questions about the future of human mathematicians and the value of human understanding in mathematics
Artificial intelligence is rapidly altering the landscape of mathematical research, with recent developments in large language models (LLMs) and automated proof assistants enabling software to generate, formalize, and verify mathematical proofs at a level previously reserved for highly trained specialists. These advances have prompted mathematicians and computer scientists to reconsider the purpose and practice of mathematics in an era where AI can perform tasks once thought to require uniquely human insight.
AI in Mathematical Proof and Discovery
For decades, computers have played a supporting role in mathematics, accelerating calculations and checking cases that would be infeasible for humans. The use of computers to prove the four-color theorem in the 1970s marked a turning point, but human mathematicians remained central to formulating conjectures, devising proof strategies, and verifying results. In recent years, however, LLMs such as those developed by Google DeepMind and OpenAI have demonstrated the ability to solve advanced mathematical problems, including those featured in the International Mathematical Olympiad, and to generate publishable research-level results in areas like arithmetic geometry.
One notable milestone occurred when Google DeepMind's experimental system, Aletheia, autonomously produced research results that met the standards for publication in peer-reviewed mathematics. In another case, an OpenAI system disproved a longstanding conjecture in combinatorial geometry, a result that leading mathematicians described as significant if achieved by a human. These systems combine LLMs with proof assistants-specialized software such as Lean, Isabelle, and Rocq-that check the logical validity of proofs step by step. Traditionally, translating informal mathematical arguments into the formal language required by proof assistants has been a labor-intensive process, but LLMs are now automating much of this work, enabling faster and more reliable verification.
Collaboration and Competition
The integration of AI into mathematical research has sparked debate within the mathematics community. Some mathematicians welcome AI as a tool that can accelerate discovery and help resolve longstanding open problems, while others express concern that automation could diminish the value of human intuition and understanding. At the 12th Heidelberg Laureate Forum in 2025, discussions focused on the possibility that AI could eventually surpass human mathematicians in generating, proving, and verifying new results, potentially relegating humans to the role of interpreters or curators of machine-generated mathematics.
Despite these concerns, most mathematicians do not expect AI to fully replace human researchers. Instead, two main approaches are emerging: one that treats AI as a tool to enhance human understanding, and another that envisions collaborative teams of humans and AI systems tackling problems together. In both scenarios, the ability to formalize and verify proofs using proof assistants is seen as a way to ensure rigor and transparency, allowing contributions from a wider range of participants-including those whose work might otherwise be overlooked.
Risks and Limitations
The growing role of AI in mathematics raises several risks and open questions. Access to advanced AI tools may become concentrated among well-funded institutions, potentially making mathematical research less accessible to individuals and smaller organizations. There is also concern that reliance on AI could erode the motivation for deep, independent engagement with difficult problems, especially among students and early-career researchers. If AI systems routinely provide answers without requiring human struggle, the development of mathematical intuition and creativity could be undermined.
To address these challenges, mathematicians and institutions are developing guidelines for the responsible use of AI in research and publication. Efforts include organizing workshops, publishing essays, and establishing standards for transparency and verification. These initiatives reflect a broader attempt to retain human agency and oversight in the evolving practice of mathematics, even as AI systems become increasingly capable.
Numerical Context
Recent AI systems have demonstrated the ability to solve problems at the level of the International Mathematical Olympiad, which features six challenging problems per year and attracts the world's most talented high school mathematicians. In 2024, Google DeepMind's Aletheia produced research-level results in arithmetic geometry, while OpenAI's system disproved a conjecture that had stood for decades. Proof assistants such as Lean and Isabelle have been used to formalize complex proofs, including the 8-dimensional and 24-dimensional sphere-packing problems, with AI-assisted formalization reducing the time required from months or years to days or weeks. These achievements highlight both the technical progress and the scale of change underway in mathematical research.
Understanding the distinction between human and AI-generated proofs requires familiarity with formal verification. Proof assistants are software systems that check each logical step in a mathematical argument, ensuring that no assumptions are omitted and that every inference is valid. While human mathematicians often rely on shared intuition and skip steps considered obvious within the community, formal proofs require complete explicitness. The automation of this process by AI not only increases reliability but also changes the nature of mathematical collaboration, making it possible for contributions to be evaluated on their formal merits rather than reputation or institutional affiliation.