• 5 mins read
  • Published

ChatGPT-Assisted Proof Solves Crouzeix's Conjecture After Decades

Noel Sharkey Technology, AI and robotics editor Science.Report

Post by Noel Sharkey

ChatGPT-Assisted Proof Solves Crouzeix's Conjecture After Decades Science.Report © science.report
ChatGPT-Assisted Proof Solves Crouzeix's Conjecture After Decades © science.report

A postdoctoral researcher in Beijing has used ChatGPT to help resolve Crouzeix's conjecture, a matrix problem that challenged mathematicians for over 20 years, raising new questions about AI's role in mathematical discovery

Jin Shanmu, a postdoctoral researcher at Peking Union Medical College in Beijing, has produced a proof of Crouzeix's conjecture-a longstanding open problem in matrix analysis-by leveraging the capabilities of ChatGPT, OpenAI's large language model. The result, which has not yet undergone peer review, has been reviewed by several mathematicians familiar with the problem, who have provisionally confirmed its correctness. This development highlights the expanding role of generative AI in mathematical research, particularly in domains where formal reasoning and symbolic manipulation are central.

Crouzeix's conjecture, first proposed by French mathematician Michel Crouzeix in the early 2000s, concerns the behavior of functions applied to matrices. Specifically, it posits that for any square matrix and any function analytic on its numerical range, the norm of the function applied to the matrix does not exceed twice the maximum of the function's modulus on that range. Despite its abstract formulation, the conjecture has implications for numerical linear algebra, operator theory, and computational physics. For over two decades, the conjecture resisted proof, with partial results and computational evidence accumulating but no general solution emerging.

AI in Mathematical Reasoning

Jin's approach stands out not only for its outcome but for the method: he used ChatGPT as a reasoning assistant, prompting the model to explore possible proof strategies and check logical steps. Unlike traditional mathematical software, large language models like ChatGPT are trained on vast corpora of text, including mathematical literature, but do not perform symbolic computation in the manner of computer algebra systems. Instead, they generate plausible next steps based on statistical patterns in language and mathematical argumentation. Jin, whose formal mathematical training is limited-his undergraduate degree is in geology and his doctoral work is in medicine-used the model to supplement his own reasoning, iteratively refining arguments and identifying gaps.

The process illustrates both the promise and the limitations of current generative AI in mathematics. While ChatGPT can suggest lines of reasoning and recall relevant theorems, it does not guarantee correctness or completeness. Human oversight remains essential, both to interpret the model's suggestions and to verify the logical validity of each step. In this case, Jin's manuscript was reviewed by mathematicians including Alex Townsend of Cornell University, Anne Greenbaum of the University of Washington, and Michel Crouzeix himself, who found the proof to be correct. However, the result awaits formal peer review and broader scrutiny by the mathematical community.

Verification and Broader Context

The announcement follows a period of increasing interest in the use of large language models for mathematical problem-solving. In May, OpenAI reported that an internal reasoning model had solved the planar unit distance problem, another longstanding mathematical challenge. The company has also listed several other problems where its upcoming model, Astra, has made significant progress. Competing AI developer Anthropic has stated that its Claude model is being used to attempt the Riemann hypothesis, one of the most famous unsolved problems in mathematics. These developments suggest that AI systems are becoming increasingly capable of contributing to mathematical research, though their outputs require careful human evaluation.

Despite these advances, the use of generative AI in mathematics raises important questions about reliability, reproducibility, and the nature of mathematical understanding. Language models can generate convincing but incorrect arguments, and their lack of explicit symbolic reasoning means that errors may be subtle or difficult to detect. The current case demonstrates that, with expert human oversight, AI-generated reasoning can assist in solving complex problems, but it does not replace the need for rigorous verification and peer review. The broader implications for mathematical practice, education, and the publication process remain to be seen.

According to available information, Jin's proof was submitted as a preprint and reviewed informally by at least three independent mathematicians. No formal error rate or benchmark evaluation applies in this context, but the review process included manual checking of each logical step. The use of ChatGPT was limited to generating and refining arguments, not to automated theorem proving or symbolic computation. The result is notable for the absence of specialized mathematical training on Jin's part, underscoring the potential for AI tools to lower barriers to entry in technical fields-while also highlighting the ongoing need for expert validation.

Large language models such as ChatGPT are trained to predict the next word or token in a sequence, based on patterns learned from massive datasets that include mathematical texts, research papers, and online discussions. While these models can generate plausible mathematical arguments and recall relevant results, they do not perform formal symbolic manipulation or guarantee logical validity. Their outputs must be interpreted and checked by human experts, especially in domains where precision and rigor are essential. The use of AI in mathematical research is likely to expand, but its reliability will depend on transparent evaluation, careful oversight, and a clear understanding of the distinction between statistical plausibility and mathematical proof.

Related articles