Anthropic's Claude tried to solve the Riemann hypothesis and found something new instead

Skye Jacobs

Posts: 2,239   +63
Staff
The big picture: Prime numbers do not follow a regular pattern, but mathematicians can measure how closely their distribution follows a predictable average. The Riemann zeta function, an infinite mathematical expression, is a key tool in that work because its nontrivial zeros contain information about the distribution of primes. The Riemann hypothesis holds that every one of those zeros lies on a single line in the complex plane, where the real part is 1/2. No one has proved it, including Anthropic's Claude. But in trying to solve the problem, Claude produced a new mathematical result that could advance the study of prime numbers.

The company says an unreleased research version of the model found a new way to show that at least 67.2% of those zeros lie on the critical line, up from the previous lower bound of 41.6%. This is a step forward on a related problem, not a solution to the Riemann hypothesis itself. That distinction is important. Mathematicians have long known that some zeros lie on the critical line; the challenge is to prove that all of them do.

According to Anthropic, Claude began by trying to tackle the full hypothesis. It did not get there. Instead, it turned to the question of how many zeros can be shown to satisfy the condition, producing the higher lower bound.

The research process also offers a look at how Anthropic is using large language models for open-ended mathematical work. Claude tested 650 ideas that did not pan out, the company said. It also used 60 agentic Claude instances to examine different approaches and check the work.

Human researchers did not provide detailed mathematical guidance, according to Anthropic. Their input largely consisted of encouragement, including telling the model to "believe in itself" and "keep going."

That is different from asking a model to solve a set problem with a known answer. In this case, the system had to explore a difficult area of number theory, discard weak approaches, and find a path that could withstand mathematical scrutiny.

James Maynard, a mathematician at the University of Oxford, said the result appears to make a real contribution. "The problem was in need of a new real idea, which this new result seems to provide," Maynard told Scientific American. "It seems that the AI has made a genuinely interesting mathematical contribution."

The Riemann hypothesis has remained unsolved for nearly 170 years and carries a $1 million prize announced in 2000. But Claude's result does not put the prize within reach. Showing that a higher percentage of zeros lie on the critical line does not prove the hypothesis, because the goal is to establish that every zero lies there.

That may sound counterintuitive. The issue is that mathematicians are dealing with an infinite set of zeros. Verifying that an increasingly large number of zeros lie on the critical line cannot rule out the possibility of exceptions farther along the sequence. Those exceptions could exist without changing the results of any finite calculation.

"Even being very optimistic, there is no pathway for any of these approaches to deal with the actual Riemann hypothesis," Maynard says.

Still, the work has drawn attention because it suggests AI systems may be useful in mathematical research beyond established benchmarks. Andrew Sutherland, a mathematician at the Massachusetts Institute of Technology, said the result points to a broader role for AI in the field.

"This is further evidence that AI is capable of doing interesting mathematical research, as opposed to just solving specific problems that are fed into it," Sutherland says.

Maynard also praised Anthropic's presentation of the research. AI companies have faced criticism when their systems reproduced existing mathematical work or presented earlier human results as original findings. In this case, he said, Anthropic gave appropriate credit to prior work and did not overstate what Claude had accomplished.

The blog post, he says, was "remarkably restrained in avoiding overhyping and giving due credit to various previous works."

Permalink to story:

 
A proof will have to be found that a nontrivial zero lying outside the critical line is an absurdity.

Penrose holds that mathematical insight requires "intuition," and that it comes to mathematicians inexplicably, usually after wrestling with a problem. He maintains that algorithmic methods can never construct a proof or discover new regions of maths. If these machine-learning models end up being able to do this, it would bring up interesting questions about the nature of their and our minds.
 
A proof will have to be found that a nontrivial zero lying outside the critical line is an absurdity.

Penrose holds that mathematical insight requires "intuition," and that it comes to mathematicians inexplicably, usually after wrestling with a problem. He maintains that algorithmic methods can never construct a proof or discover new regions of maths. If these machine-learning models end up being able to do this, it would bring up interesting questions about the nature of their and our minds.

Dam you, you actually made me use my brain for a few minutes. Seriously though, that’s a genuinely interesting angle. If AI starts producing original mathematical proofs consistently, the old distinction between human ‘intuition’ and machine computation gets a lot harder to define. Very good post!
 
A proof will have to be found that a nontrivial zero lying outside the critical line is an absurdity.

Penrose holds that mathematical insight requires "intuition," and that it comes to mathematicians inexplicably, usually after wrestling with a problem. He maintains that algorithmic methods can never construct a proof or discover new regions of maths. If these machine-learning models end up being able to do this, it would bring up interesting questions about the nature of their and our minds.
Penrose is an interesting guy, but clearly a bit out there. Clearly math can be mathed without intuition. People discover things on accident all the time.
 
Dam you, you actually made me use my brain for a few minutes. Seriously though, that’s a genuinely interesting angle. If AI starts producing original mathematical proofs consistently, the old distinction between human ‘intuition’ and machine computation gets a lot harder to define. Very good post!
Must have been painful.
 
Penrose is an interesting guy, but clearly a bit out there. Clearly math can be mathed without intuition. People discover things on accident all the time.
“Clearly math can be mathed without intuition” is certainly a sentence, but it rather skips over the entire philosophical question GeoffreyA raised.

Accidental discovery doesn’t eliminate intuition. Plenty of things are encountered accidentally...the difficult part is recognizing that the anomaly is meaningful, understanding what it implies, connecting it to existing structure, and then proving that the relationship is not merely coincidence.

Penrose’s argument is not “mathematicians never stumble onto things.” It is about whether mathematical understanding can be reduced entirely to formal, algorithmic procedure.

Which is precisely why this result is interesting.

If a machine can search an enormous mathematical space, identify a non obvious relationship, distinguish a productive direction from millions of useless ones, assemble existing ideas in a novel way, and produce a verifiable proof, then the interesting question becomes whether we have simulated something resembling intuition...or whether what we romantically call “intuition” was computational all along.

But yes, I suppose “people find stuff by accident” settles several decades of philosophy of mathematics rather efficiently.

Someone should probably let Penrose know.
 
“Clearly math can be mathed without intuition” is certainly a sentence, but it rather skips over the entire philosophical question GeoffreyA raised.

Accidental discovery doesn’t eliminate intuition. Plenty of things are encountered accidentally...the difficult part is recognizing that the anomaly is meaningful, understanding what it implies, connecting it to existing structure, and then proving that the relationship is not merely coincidence.

Penrose’s argument is not “mathematicians never stumble onto things.” It is about whether mathematical understanding can be reduced entirely to formal, algorithmic procedure.

Which is precisely why this result is interesting.

If a machine can search an enormous mathematical space, identify a non obvious relationship, distinguish a productive direction from millions of useless ones, assemble existing ideas in a novel way, and produce a verifiable proof, then the interesting question becomes whether we have simulated something resembling intuition...or whether what we romantically call “intuition” was computational all along.

But yes, I suppose “people find stuff by accident” settles several decades of philosophy of mathematics rather efficiently.

Someone should probably let Penrose know.
It's not one or the other. Penrose is drawing a line in the sand because he's old, and he wants math to be done by humans. I don't disagree with his feelings, it's simple not logical; at all.
 
It's not one or the other. Penrose is drawing a line in the sand because he's old, and he wants math to be done by humans. I don't disagree with his feelings, it's simple not logical; at all.
That’s an interesting way to dispose of a fairly serious philosophical argument...attribute it to Penrose being old and emotionally attached to humans, then declare it “not logical.”

Penrose’s position may ultimately be wrong, but it isn’t “humans should do math because I like humans.” His argument concerns whether mathematical understanding and insight can be fully captured by formal algorithmic processes, with Gödelian incompleteness playing a central role.

You can dispute that reasoning...and plenty of philosophers and mathematicians do, but “he’s old” is not generally considered a devastating counterexample in mathematical logic.

The Claude result is interesting precisely because it gives us something empirical to test against Penrose’s position. If machines increasingly produce genuinely novel mathematics rather than merely executing predefined procedures, then we get to examine whether the distinction he drew actually survives.
 
Penrose’s argument is not “mathematicians never stumble onto things.” It is about whether mathematical understanding can be reduced entirely to formal, algorithmic procedure.
Yes. If I remember correctly, Penrose argued that a mathematical truth can't be verified with a proof according to the rules of that system. It had to do with Gödel's incompleteness theorem. Therefore, an algorithm, working within the confines of a formal mathematical system, could not construct a valid proof. When mathematicians did so, it was by way of perception outside the system. (Ultimately, this may go into the metaphysical, asking where do maths systems "exist" and Platonism, reaching to the very bedrock of reality, dizzying heights for the human mind.)

If a machine can search an enormous mathematical space, identify a non obvious relationship, distinguish a productive direction from millions of useless ones, assemble existing ideas in a novel way, and produce a verifiable proof, then the interesting question becomes whether we have simulated something resembling intuition...or whether what we romantically call “intuition” was computational all along.
That's the point which always strikes me. Is our intuition or inspiration "merely" a computational system? If so, it can be matched, in principle, by the machine.

(Personally, I think that's the case. Perhaps we have a "maths model" working in concert with an "imagination model," and in the cases of certain geniuses, Newton, Euler, Penrose, etc., it creates what we call mathematical insight, leading to new systems, proofs, etc.)

Penrose’s position may ultimately be wrong, but it isn’t “humans should do math because I like humans.” His argument concerns whether mathematical understanding and insight can be fully captured by formal algorithmic processes, with Gödelian incompleteness playing a central role.

You can dispute that reasoning...and plenty of philosophers and mathematicians do, but “he’s old” is not generally considered a devastating counterexample in mathematical logic.

The Claude result is interesting precisely because it gives us something empirical to test against Penrose’s position. If machines increasingly produce genuinely novel mathematics rather than merely executing predefined procedures, then we get to examine whether the distinction he drew actually survives.
You summed it up better than I could.
 
Last edited:
Penrose is an interesting guy, but clearly a bit out there. Clearly math can be mathed without intuition. People discover things on accident all the time.
With the disclaimer that he's my favourite living physicist, I agree that he can be idiosyncratic, but hence comes the strength and originality of his thought, and a boldness to question, what he would call, scientific dogma or fashion, whether it is the weaknesses of quantum mechanics, inflationary, or string theory---the Susskind camp in other words! On some points, I do disagree with him, such as his old belief that consciousness relies on the quantum.
 
Whilst I accept that these LLMs aren't probably able to produce actual proofs (yet), the success they seem to be having in mathematics seems to suggest most mathematicians have been coasting most of their careers... 😁
 
If a machine can search an enormous mathematical space, identify a non obvious relationship, distinguish a productive direction from millions of useless ones, assemble existing ideas in a novel way, and produce a verifiable proof, then the interesting question becomes whether we have simulated something resembling intuition...or whether what we romantically call “intuition” was computational all along.

Mmm, I'm not convinced. Intuition of other researchers has been captured in scientific literature, public communications of scientists, etc. So it has become a documented path. Second, a machine can brute force things. After all, it did try 650 different paths already as stated in the article.
 
Mmm, I'm not convinced. Intuition of other researchers has been captured in scientific literature, public communications of scientists, etc. So it has become a documented path. Second, a machine can brute force things. After all, it did try 650 different paths already as stated in the article.
I think that’s a fair objection, but I’m not sure it actually resolves the intuition question.

Yes, Claude is drawing from human mathematics already captured in papers, lectures, proofs, heuristics, and decades of accumulated research. But human mathematicians do exactly that too. Nobody develops mathematical intuition in a vacuum...it is built from prior knowledge, patterns, methods, and experience.

Trying 650 paths certainly shows a lot of search, but brute force by itself is not enough. The harder part is deciding which paths are worth pursuing, when to abandon one, which intermediate result is significant, and when two seemingly unrelated ideas might fit together.

That selection process is precisely where the word “intuition” starts becoming interesting.

So I’d agree that Claude may be leveraging human intuition encoded in the literature plus enormous computational search. What I’m not convinced of is that this somehow makes the result philosophically uninteresting. If anything, it raises the question of whether human intuition is also largely learned pattern recognition, search, and selection operating below conscious awareness.
 
If anything, it raises the question of whether human intuition is also largely learned pattern recognition, search, and selection operating below conscious awareness.

Personally, I'm pretty sure that it's (mostly) based on what you say.

Without detailed insight in how Claude functions, we cannot say anything on why certain paths where given up. Maybe the thinking budget per path was limited, maybe there were no more ends to follow, maybe there is some sort of scoring system at work, maybe... and maybe it can even be called intuition.
 
Personally, I'm pretty sure that it's (mostly) based on what you say.

Without detailed insight in how Claude functions, we cannot say anything on why certain paths where given up. Maybe the thinking budget per path was limited, maybe there were no more ends to follow, maybe there is some sort of scoring system at work, maybe... and maybe it can even be called intuition.
I think that’s a fair way to put it, and I’m pretty close to your position.

Without visibility into Claude’s internal process, we really can’t say whether a path was abandoned because of a hard budget, a scoring mechanism, learned heuristics, lack of promising continuations, or something more analogous to what we’d call intuition.

And that may be the most interesting part of the whole issue. We don’t have especially good visibility into our own intuition as well. We experience the result of a lot of hidden mental processing and then give it a name.

Me? I’m not ready to say Claude has intuition, but I’m also not convinced we currently understand intuition well enough to confidently say it doesn’t.
 
Back