Why It Matters
The release of these 722 mathematical papers highlights the growing tension between proprietary AI development and the open-source movement, raising questions about transparency and collaboration in the field. By selectively publishing outputs from an undisclosed model, OpenAI may be setting a precedent for how AI-generated research is disseminated, potentially influencing regulatory discussions around intellectual property and ethical considerations in AI usage. This situation invites scrutiny from both the academic community and industry stakeholders who are increasingly concerned about the implications of opaque AI systems on innovation and trust.
OpenAI put 722 mathematical manuscripts on GitHub last week. All of them came from an AI model the company hasn’t named, hasn’t released, and won’t let outsiders poke around in. That’s basically the whole controversy in two sentences.
The papers were organized into 372 “families” — OpenAI’s word for clusters of related results. The company fed roughly 4,000 problems to the AI, then hand-picked whatever it considered significant enough to publish. Nearly all outputs came from a single prompt given to one AI agent, though some problems needed multiple attempts before the model produced anything worth keeping. What that selection process actually looked like, who made the calls, and what got thrown out — unclear.
Only 162 papers checked out.
That’s the number with computer-verified main results using Lean, a software tool that checks whether logical steps in a proof actually hold. Out of 722 papers, 162 have that stamp. That’s roughly 22%. OpenAI acknowledged, pretty directly, that some of the unformalized results might contain errors. Which is a strange thing to publish, but here we are.
What the Academic World Actually Thinks
Andrew Sutherland from MIT didn’t exactly celebrate. Without the model being available for public verification, he said, the results should be treated as unverified. That’s a careful, measured reaction — but in academic terms, it’s a cold bucket of water. Verification is kind of the whole point of mathematics. A result that can’t be checked isn’t really a result yet.
Daniel Litt from the University of Toronto went further, arguing against keeping mathematical answers confidential at all. The logic there is straightforward: math is a shared enterprise. Locking up the process — the model, the prompts, the reasoning chain — cuts everyone else out of the work.
The Institute for Advanced Study in Princeton weighed in too. The institute’s advisory group had previously called on OpenAI to release more details: the model name, the prompts used, the full methodology. OpenAI provided some average compute figures and reasoning summaries. But the full transparency the advisory group asked for? Not there yet. The institute was clear that mathematical understanding needs to stay central, even as AI tools get more capable.
OpenAI also turned off Issues on its GitHub repository, which means outside researchers can’t file bug reports, flag errors, or push back through the platform directly. That’s a pretty significant move for something positioned as a contribution to the broader mathematical community. It limits the feedback loop that normally helps science self-correct.
The Anthropic Comparison Nobody’s Ignoring
Anthropic did something different. Its team produced a Lean-checked proof of Fermat’s Last Theorem and made all the data publicly available. The proof itself wasn’t a new discovery — it revalidated a theorem Andrew Wiles originally proved in 1995 — but the process was open. Anyone could look at the proof, check the Lean formalization, and see exactly how the result was reached.
OpenAI’s 722 manuscripts sit in a different category. The company describes many of them as “exceptional advances within an existing program” rather than transformative breakthroughs. One problem that drew particular attention was the Quasi-Riemann Hypothesis. But the framing throughout is careful — these aren’t millennium problems. They’re significant, probably, but the full picture won’t be clear until more of the work is formalized and checked.
OpenAI says it plans to keep adding Lean formalizations over time. So the 22% figure isn’t meant to be the final number. But “we’ll do more later” isn’t the same as “here’s the model, here’s the process, verify it yourself.” And that gap is what’s driving most of the frustration.
There’s a real tension sitting underneath all of this. AI-generated mathematics is genuinely new territory. The tools are getting fast enough and capable enough that a single model, given a single prompt, can produce hundreds of papers in a run. That’s wild. It’s also exactly why transparency matters more, not less. The faster the output, the harder it is to catch errors without open access to the process.
Some researchers see the release as a real milestone — evidence that AI can engage with serious mathematical problems at a level that wasn’t possible a few years ago. Others are more cautious, and not without reason. A paper that might contain errors, produced by a model nobody can examine, published in a repository that doesn’t allow outside contributions, is a hard thing to evaluate. The excitement is real. So is the skepticism.
OpenAI’s commitment to adding more formalizations over time is at least something. The academic community is watching. What they’re watching for, mostly, is whether the company eventually opens the model up — or whether the 722 manuscripts stay the property of a process nobody outside OpenAI can fully see.
The Quasi-Riemann Hypothesis paper is still sitting there, unformalized.
Frequently Asked Questions
What did OpenAI release on GitHub?
OpenAI published 722 mathematical manuscripts generated by an undisclosed AI model, organized into 372 families of related results, selected from roughly 4,000 problems the AI was given.
How many of the 722 papers have verified results?
Only 162 of the 722 papers include computer-checked main results using Lean, which works out to about 22% of the total collection.
Why are mathematicians pushing back on the release?
Researchers including Andrew Sutherland of MIT and Daniel Litt of the University of Toronto say the results can’t be properly verified without access to the model itself, and OpenAI’s decision to disable external contributions on GitHub limits the academic community’s ability to engage with the work.
