On August 1, OpenAI announced that it had produced new results on ten long-standing open problems in mathematics and theoretical computer science. The work used an internal version of Astra, a next-generation model family that has not yet been released to the public.

The problems span eight areas, including high-dimensional geometry, coding theory, group theory, quantum complexity, lattice-based cryptography, and extremal combinatorics. The ten results are not all of the same kind. Some provide new solutions, some produce counterexamples to existing conjectures, and others push beyond previously known bounds. It is therefore more accurate to describe the announcement as a mix of solved problems and meaningful advances rather than saying that AI simply “solved ten hard problems.”
Machine Verification of AI-Generated Proofs
To understand why this announcement matters, it is important to look not only at the results themselves, but also at how OpenAI chose to verify and release them. When an AI claims to have answered an open problem, an obvious question follows: how can we tell whether the answer is actually correct—and whether it is genuinely new?
OpenAI had already faced a similar controversy. In October 2025, Kevin Weil, who was then leading OpenAI’s science research initiative, OpenAI for Science, wrote that GPT-5 had solved ten open problems posed by mathematician Paul Erdős and made progress on eleven more. But the situation changed after Thomas Bloom, who maintains the Erdős Problems database, and other mathematicians examined the claims. What the model had found were not new solutions. They were existing solutions that had already been published by other researchers but had not yet been reflected in the database. In other words, GPT-5 had largely found answers that already existed outside the database rather than creating new ones. Weil later acknowledged that he had misinterpreted the result and deleted the post.

The episode showed why it is important to distinguish between “search” and “discovery” when AI enters the research process. Finding an answer that already exists somewhere in the literature is fundamentally different from producing a solution that no one knew before. For the new set of ten results, simply saying that AI had found an answer was therefore not enough. There also needed to be a way for outside researchers to check whether the result actually held.
That is where Lean 4 comes in. Alongside roughly 250 pages of research results, OpenAI released Lean 4 formalizations of all ten results on GitHub. The repository includes Lean files corresponding to results on sphere packing, non-sofic groups, Connes’ rigidity conjecture, quantum parallel repetition, and other problems. Outside researchers can run the files themselves and check whether the proofs pass Lean’s formal rules.

Lean is not another AI that reads a proof and says, “This looks correct.” It is a proof assistant that translates mathematical definitions and logical steps into a form a computer can process, then checks whether each conclusion actually follows from the preceding assumptions under strict formal rules. In an ordinary mathematical proof, a researcher may skip familiar steps or move past something that seems obvious. In Lean, the necessary logical steps have to be made explicit. If the next step does not follow from the rules, the proof does not pass.
A programming analogy makes this easier to understand. It is closer to compiling code and checking whether it produces errors than to having someone read the code and say that it looks reasonable. Mathematical proofs and software are not the same thing, of course, but the similarity lies in the fact that a long argument no longer has to be checked only through human intuition.
A mathematical argument translated into a strict form that Lean can check is called a “formal proof.” At that stage, a computer can verify whether an already constructed line of reasoning is logically consistent. But deciding how to approach a problem in the first place, or how to combine ideas from different areas, is a separate task. When OpenAI says that Astra generated the argument, it is referring to this earlier part of the process.
According to OpenAI, Astra first generated the mathematical argument, researchers then organized it into a paper, and the model was used again to formalize the argument as a Lean proof. In other words, the natural-language proof was not only read and assessed by humans; it was also translated into a form that a computer could check. As a result, the announcement goes beyond simply saying, “The AI produced an answer.” Other researchers can use the released Lean proofs to check directly whether the formalized arguments satisfy Lean’s logical rules.

Lean, however, cannot make every judgment involved in research. It can check whether the logic follows correctly within the formalized statement and assumptions, but it cannot automatically determine whether the formalized problem precisely matches the original open problem. Nor can it decide whether a similar method has already appeared in another paper, how original the result is, or how important it should be considered.
This creates two different kinds of verification. One asks whether the proof is logically valid. The other asks whether the result is genuinely new and meaningful as research. Machines can now play a substantial role in the first. The second still depends heavily on human judgment.
The Debate Over Novelty and Credit
Once AI began producing answers to actual research problems, another question followed naturally. Is a result that is logically valid necessarily a new research contribution?
Much of the criticism that followed OpenAI’s announcement focused on this distinction. Interestingly, the main argument was not that the proofs were wrong. Mathematicians instead questioned how heavily some of the results depended on recent prior work, and whether those earlier contributions had been represented clearly enough in the announcement.
A representative example is the high-dimensional sphere-packing problem. Sphere packing asks how densely equal-sized spheres can be arranged in space without overlapping, and the problem becomes much more complicated in higher dimensions.

Astra improved an existing upper bound on how densely spheres can be packed. But in an interview with Scientific American, Yeshiva University mathematician Stephen Miller argued that one of the central arguments in the proof was effectively the same as a method he and his collaborators had used in a 2016 paper. He claimed that OpenAI had not given sufficient recognition to that prior work and went as far as raising concerns about research ethics. This remains Miller’s allegation; it does not mean that an official finding of research misconduct has been made.
The debate over non-sofic groups went a step further. While the sphere-packing controversy was largely about whether prior work had been properly acknowledged, the non-sofic-group case raised a deeper question: where should the line be drawn between earlier human research and the AI’s own new contribution?
In mathematics, a “group” is a structure that organizes transformations—such as symmetries or rotations—according to rules that allow them to be combined. In group theory, mathematicians had long asked whether every group has a property known as “sofic,” or whether groups without that property actually exist. Astra provided a concrete example of a non-sofic group, offering an answer to that long-standing question.
But Francesco Fournier-Facio of the University of Cambridge and other researchers in the field argued that the result did not begin from an entirely new starting point. In particular, work published by Andreas Thom and Gábor Kun in 2016 and 2019 had already provided important foundations for the final step toward the construction.
After these criticisms, OpenAI revised the wording of its official announcement. The original version had described the target problems as having seen no major progress for at least a decade. The current version instead says that the results “resolve or make substantial progress on long-standing open problems.” OpenAI also acknowledged that the non-sofic-group proof relied heavily on significant recent research and adjusted its description to reflect those prior contributions more accurately.
The controversy did not amount to a rejection of Astra’s result itself. The real question was how far earlier researchers had already paved the way, and what exactly Astra had added on top of that foundation. Even if Kun’s work had already advanced much of the problem, identifying the remaining connection and turning it into an explicit construction of a non-sofic group can still be considered a distinct contribution. The debate was therefore less about whether there was a result at all, and more about how much of the final outcome should be regarded as a new discovery.
This problem of credit did not begin with AI. Mathematics has always advanced by building on earlier methods and results. One researcher may develop a key technique, another may improve it, and a third may complete the final step. The person who delivers the final solution often receives the most attention, even though the result was only possible because of the work that came before.
AI makes this old problem of attribution more visible. Fournier-Facio compared mathematical research to a pyramid built over generations. The person who places the final stone is the most visible, but many layers of work already exist underneath. If the final stone is placed by an AI, there is a risk that the result itself will receive attention while the human research beneath it becomes less visible.
This is also where the limits of Lean become clear again. Lean can check whether a formalized proof follows logically. It cannot determine whether the same idea appeared in another paper years earlier, how much an existing method must change before it counts as new research, or which researcher’s contribution was decisive.
A logically correct result and an academically novel result are not the same thing.
OpenAI’s announcement showed that AI can participate in mathematical research, but it also showed that the scope of verification has expanded. It is no longer enough to ask whether a proof is correct. Researchers also have to ask where the result came from and what earlier work it depends on.
As AI produces research answers more quickly, those judgments may become even more important. And that leads to the next question: what role will researchers themselves play in this changing process?
How the Research Process May Change in the AI Era
The most immediate change suggested by this announcement is not that AI is about to replace researchers. A more realistic shift is that the tasks performed by humans and machines within the research process are beginning to separate.
AI is moving beyond searching and summarizing existing material and into the generation of candidate solutions to genuine open problems. Tools such as Lean allow part of the verification process to be delegated to machines. But deciding which problems matter, whether a result is truly new, and how it relates to previous work still requires human judgment.
If AI can generate more candidate solutions at much greater speed, the role of the researcher may change as well. Rather than carrying out every step alone, researchers may spend more time deciding which questions deserve attention, identifying which results are worth pursuing, and turning one answer into the next research problem.
Mathematics does not advance simply by solving one problem after another. Strong results create new methods, and those methods generate new questions. It is still unclear whether today’s AI systems can independently reach that stage. The latest results clearly matter because they address genuine open problems, but whether they can also open entirely new research directions remains to be seen.
So the question left by this announcement is slightly different from “Will AI replace researchers?” It is closer to this: as producing answers becomes easier, what will humans need to become better at?
Choosing worthwhile problems, judging the novelty and importance of results, and making clear where those results came from. In an era when AI can increasingly produce answers, perhaps the defining role of the human researcher will be deciding what to trust and what is worth valuing.

