TL;DR
Anthropic's research model uses autonomous agents and Lean formalization to make a breakthrough on the century-old Riemann Hypothesis math problem.
A research version of Claude has significantly advanced the search for a solution to the Riemann Hypothesis, a problem that has eluded mathematicians since 1859. The model successfully increased the longstanding mathematical lower bound proportion of zeros on the critical line for the Riemann zeta function from 41.6% to 67.2%.
This progress occurred during an autonomous multi-day testing session. Unlike standard chatbot interactions, the model operated within Claude Code, consuming 31 million output tokens across two distinct sessions to iterate on the proof.
The path to the result was not linear. The system generated 650 failed hypotheses before coordinating 60 subagents to execute 2,400 shell commands and write hundreds of Python scripts. This agentic approach allowed the model to explore a vast search space of mathematical possibilities that would be prohibitively slow for a human researcher.
Technical validation
To ensure the result was not a hallucination, Anthropic employed a rigorous verification pipeline. Internal mathematicians Levent Alpöge and Ralph Furman first validated the findings, followed by external review from number theorists Brian Conrey and Dan Goldston. Finally, staff member Eric Easley formalized the proof using Lean, a formal verification language that provides a machine-checked guarantee of correctness.
According to neowin.net, the model did not invent a new branch of mathematics from scratch. Instead, it synthesized and combined existing research frameworks from several prominent mathematicians, including Bombieri, Goldston, and Turnage-Butterbaugh, to find the new bound.
This capability represents a shift in how artificial intelligence handles high-level reasoning. While current public models often provide generic explanations when asked about the Millennium Prize Problems, this research version demonstrated the ability to perform genuine discovery through autonomous iteration and tool use.
Broader AI landscape
The breakthrough arrives amid a surge of specialized model releases. While Anthropic pushes into theoretical mathematics, Meta has focused on accessibility with the open-source Muse Glimmer, designed for local hardware as reported by theamericanconservative.com. Meanwhile, OpenAI is pivoting toward defensive security with its Daybreak initiative and the GPT-5.6-Cyber model, as detailed by cnbc.com.
These parallel developments suggest a fragmentation of frontier AI. We are seeing a transition from general-purpose assistants to highly specialized agents capable of breaking through specific scientific bottlenecks or securing infrastructure. The use of multi-agent orchestration in the Riemann case mirrors the architecture of Nvidia's Nemotron 3 family, which humanityredefined.com notes is specifically optimized for collaborative, long-task workflows.
For practitioners, the most critical takeaway is the integration of LLMs with formal verification tools like Lean. The ability of a model to propose a hypothesis is valuable, but the ability to formalize it into a provable statement is what transforms a stochastic guess into a mathematical fact. This loop of hypothesis, execution, and formal verification is the new blueprint for AI-led scientific discovery.
As models move from predicting the next token to solving century-old problems, the boundary between heuristic search and genuine mathematical intuition continues to blur. The question remains whether AI will eventually provide the final proof for the Riemann Hypothesis or if it will remain a powerful tool for narrowing the search space for humans.
FAQ
What is the Riemann Hypothesis?
It is a conjecture about the distribution of prime numbers, specifically regarding the zeros of the Riemann zeta function. It is one of the seven Millennium Prize Problems.
Did Claude fully solve the problem?
No. It improved a specific mathematical lower bound (from 41.6% to 67.2%), which is significant progress but not a complete proof of the hypothesis.
How was the proof verified?
It was reviewed by internal and external mathematicians and then formalized in Lean, a formal proof assistant.
Is this model available to the public?
No, this was an unreleased research version of Claude used in a controlled environment.
About the Author
Guilherme A.
Former dentist (MD) from Brazil, 41 years old, husband, and AI enthusiast. In 2020, he transitioned from a decade-long career in dentistry to pursue his passion for technology, entrepreneurship, and helping others grow.
Connect on LinkedIn