A New Round in the Argument Over What AI Can Discover
OpenAI’s latest claim — that an internal model produced 10 advances on long-stalled problems in mathematics and theoretical computer science — has electrified parts of the research world while sharpening a familiar dispute over how far today’s artificial intelligence can really go.
The company said on Friday that an internal version of its forthcoming model, called Astra, generated results spanning geometry, complexity theory, cryptography, Ramsey theory and extremal graph theory. OpenAI said the work was later written up by humans and formalized in Lean, a proof assistant used to verify mathematical arguments line by line. It also said the successful solution runs would have cost roughly $2,000 each at current API token prices.
For supporters, the announcement is another sign that frontier AI systems are beginning to do something more consequential than solving olympiad-style exercises or acing benchmarks: contributing to research-level mathematics on problems that experts had left unresolved for years. For skeptics, it is also a reminder of how quickly spectacular claims can outrun what is actually known about a system’s reliability, generality and scientific importance.
That tension has become one of the defining arguments in AI this year.
Why Math Has Become the Test Case
Mathematics occupies a special place in the broader debate over AI and discovery. Unlike claims in biology or chemistry, where results may depend on noisy experiments, hidden variables or months of laboratory follow-up, mathematical proofs can in principle be checked precisely. Formal verification tools like Lean make that even more stringent, allowing other researchers to confirm whether an argument holds.
That is one reason math has emerged as an early proving ground for research-oriented AI. OpenAI’s latest release builds on a series of similar disclosures this year, including a February project testing whether models could generate expert-checkable proofs and a May announcement that an internal system had disproved the Erdős unit-distance conjecture, a famous problem in combinatorial geometry.
By releasing papers, proof certificates and model-generated walkthroughs, OpenAI is plainly trying to frame the new results not as vague promises of scientific revolution but as auditable outputs in a domain where verification is unusually clean.
Even some critics of AI hype acknowledge that this matters. A proof that can be checked by independent mathematicians is different from a chatbot’s confident explanation or a benchmark score. If the results stand, they would add to the evidence that current systems can help produce genuine, novel mathematics — at least in some categories of problems where search, abstraction and formal reasoning can be combined effectively.
Excitement, and Unease, Among Mathematicians
The reaction from mathematicians has been strikingly mixed: admiration, curiosity and, in some quarters, dread.
Timothy Gowers, the Fields Medal-winning mathematician, wrote in late July that he had twice seen GPT-5.6 Pro produce one-shot solutions to problems he had spent serious time thinking about himself. He described the experience as both deeply impressive and unsettling. His concern was not simply whether the answers were right, but what happens to mathematical culture if researchers stop developing the habits of mind needed to discover and understand such proofs on their own.
That anxiety has been building as AI systems move from assisting with exposition and computation into areas closer to creative research. The prospect is not merely that machines will speed up mathematics, but that they could alter who does it, how credit is assigned and what kinds of expertise remain valuable.
Some researchers have embraced that possibility more readily. Terence Tao, the Fields Medalist who has written extensively about AI in mathematics, has argued that the technology could usher in an era of “big mathematics,” with human researchers and machines dividing labor across large, distributed collaborations. In that vision, people focus on taste, strategy and conceptual framing, while AI systems handle more of the technically exhaustive work.
The new OpenAI announcement has given that idea fresh momentum.
The Transparency Question
Still, the company’s release has left many unanswered questions.
OpenAI published a paper, Lean formalizations and reconstructed proof walkthroughs, which is more documentation than many AI announcements provide. But critics have pointed to what remains missing: the full prompts, the number of failed attempts, the degree of human steering and the methodological controls needed to judge whether the work reflects robust capability or a narrower, hard-to-interpret success mode.
That matters because the headline number — around $2,000 in tokens for a successful run — has circulated widely online, sometimes as if it represented the all-in cost of automating mathematical discovery. It does not. The figure excludes unsuccessful runs, model training, engineering overhead and the labor of experts who translated the outputs into manuscripts and verified the underlying arguments.
The distinction is important. One reading of the result is that AI is becoming remarkably cheap at generating candidate insights in formal domains. Another is that a highly resourced laboratory, using a closed and unreleased model plus expert oversight, has shown a still-unclear level of research assistance under unusually favorable conditions.
Both can be true at once.
What This Does — and Does Not — Say About Science More Broadly
The broader stakes lie beyond mathematics. AI companies have increasingly suggested that the same advances powering coding and reasoning systems could soon accelerate science itself, from drug discovery to materials design.
But recent research has given skeptics ammunition. Several studies from academic groups this year found that even strong frontier models struggle with the messier aspects of open-ended scientific work: generating truly novel hypotheses, designing informative experiments, iterating after failed results and distinguishing promising ideas from plausible-sounding dead ends. In one physics-discovery benchmark, top-performing agents reportedly solved only about half of simulated worlds and had particular difficulty uncovering hidden structure through experiment design.
That gap is central to the current backlash against grander claims. Many researchers see mathematics as a special case — unusually structured, richly represented in digital text and, above all, verifiable. An AI system that can help prove theorems is not necessarily one that can run a biology lab, formulate an original theory in condensed-matter physics or navigate the ambiguities of real-world data.
This is the line skeptics have pressed in recent days. They argue that the new results, while notable, should not be mistaken for evidence that automated science is imminent. Formal theorem proving, on this view, is a meaningful but domain-constrained success.
Supporters counter that this objection can become a moving target: each time systems clear a previously doubted bar — coding, formal reasoning, algorithmic insight, now research mathematics — the standard for what counts as “real” intelligence shifts again. On social media, that argument has fueled a separate backlash against prominent AI skeptics accused of minimizing each new advance after the fact.
Validation Will Take Time
For now, the immediate question is narrower and more concrete: whether the 10 claimed advances hold up under sustained scrutiny from the mathematical community.
Lean verification should help settle basic correctness questions relatively quickly where formalizations are complete. But significance is not binary. Mathematicians will still debate how surprising the results are, how much conceptual novelty they contain, whether they open new lines of inquiry and how much of the real intellectual work came from the model as opposed to the humans around it.
Those are not trivial distinctions. In science and mathematics, being right is only part of the story; understanding why a result matters is part of the achievement.
What seems increasingly clear is that AI has moved deeper into territory that many researchers once thought would remain human for longer. Even if mathematics proves to be the easiest frontier of genuine discovery, it is an important one. Proof-based disciplines offer a rare environment where claims can be tested cleanly, and where progress can be separated, at least somewhat, from marketing theater.
That is why OpenAI’s latest announcement matters now. It does not resolve the argument over machine-led science. But it makes that argument harder to dismiss — and harder to simplify.
Sources
Further reading and reporting used to add context:
- https://www.theatlantic.com/technology/2026/07/jacob-tsimerman-math-fields-medal-openai/688120/?utm_source=apple_news
- https://arxiv.org/abs/2607.09217
- https://paperswithlean.com/
- https://openai.com/research/index/publication/
- https://arxiv.org/abs/2508.15878
- https://www.reddit.com/r/OpenAI/comments/1uyeq3t/gpt_56_solved_all_6_problems_from_imo_2026/
- https://www.reddit.com/r/ChatGPT/comments/1uyerah/gpt_56_solved_all_6_problems_from_imo_2026/
- https://doi.org/10.1038/s41586-026-10652-y
- https://pubmed.ncbi.nlm.nih.gov/42420458/
- https://www.nature.com/articles/s41586-026-10265-5
- https://openai.com/index/first-proof-submissions/
- https://www.reddit.com/r/OpenAI/comments/1uw2go5/gpt56_pro_solves_5_erdos_problems/
- https://link.springer.com/article/10.1140/epjds/s13688-026-00672-z
- https://arxiv.org/abs/2606.08723
- https://news.mit.edu/2026/exposing-biases-moods-personalities-hidden-large-language-models-0219
- https://cdn.openai.com/pdf/26177a73-3b75-4828-8c91-e8f1cf27aaa0/oai_first_proof.pdf
- https://www.tomsguide.com/ai/chatgpt/chatgpt-just-solved-a-150-pokemon-crossword-with-no-clues-heres-why-thats-a-big-deal
- https://news.mit.edu/2026/teaching-ai-agents-ask-better-questions-playing-battleship-0603
- https://www.reddit.com/r/math/comments/1uydg8w/gpt_56_solved_all_6_problems_from_imo_2026/
- https://arxiv.org/abs/2605.30106
- https://www.reddit.com/r/mathematics/comments/1vcgwiu/ten_advances_in_mathematics_and_theoretical/
- https://datascience.uchicago.edu/news/2026-ai-science-hackathon-tackles-real-world-scientific-challenges-using-ai/
- https://openai.com/tr-TR/index/first-proof-submissions/
- https://openai.com/hr-HR/index/first-proof-submissions/
- https://openai.com/jv-ID/index/first-proof-submissions/
- https://openai.com/sk-SK/index/first-proof-submissions/
- https://openai.com/pl-PL/index/first-proof-submissions/
- https://openai.com/index/ten-years/
- https://openai.com/de-DE/index/first-proof-submissions/
- https://openai.com/ar/index/first-proof-submissions/
- https://openai.com/ca-ES/index/first-proof-submissions/
- https://openai.com/fil-PH/index/first-proof-submissions/
- https://openai.com/ja-JP/index/first-proof-submissions/
- https://cdn.openai.com/pdf/c4c0b11c-33cd-41e0-99aa-8f2b1a8fac2f/you-can-just-build-things-promotion-terms.pdf
- https://cdn.openai.com/pdf/045aa967-ee96-4a09-94ee-3098ddf6db2c/OpenAI-US-House-Select-Cmte-Update-%5B021226%5D.pdf
- https://cdn.openai.com/signals/openai-signals-global-report.pdf
- https://cdn.openai.com/pdf/f4b4a5da-b2de-418d-9fcd-6b293e9dc157/oai_ai-as-a-scientific-collaborator_jan-2026.pdf
- https://github.com/ten-protocol
- https://github.com/openai/openai-openapi
- https://github.com/openai
- https://github.com/OPENAI
- https://github.com/openai/model_spec
- https://github.com/orgs/openai/repositories
- https://github.com/openai/openai-python/releases
- https://github.com/openai/openai-java
- https://github.com/openai/evals/pulls
- https://cdn.openai.com/pdf/47c0215b-8976-4f60-8e13-d69c2ddbc15e/a-practical-guide-to-building-with-gpt-5.pdf
- https://github.com/openai/evals
- https://github.com/openai/miniF2F
- https://cdn.openai.com/openai_proof_assistant_tutorial.pdf
- https://github.com/openai/parameter-golf
- https://cdn.openai.com/pdf/a21c39c1-fa07-41db-9078-973a12620117/cot_controllability.pdf
- https://cdn.openai.com/improving-mathematical-reasoning-with-process-supervision/Lets_Verify_Step_by_Step.pdf
- https://theweek.com/tech/mathematicians-are-buzzing-about-an-ai-solution-to-a-math-mystery
- https://leidendeclaration.ai/
- https://arxiv.org/abs/2605.30284
- https://arxiv.org/abs/2606.08251
- https://arxiv.org/abs/2606.10587
- https://arxiv.org/abs/2601.03315
- https://the-decoder.com/openai-researchers-explain-why-math-is-the-road-to-agi/
- https://leidendeclaration.ai/signatories?page=5
- https://www.sciencedirect.com/science/article/pii/S0020025526005074
- https://the-decoder.com/terence-tao-says-gpt-5-2-pro-cracked-an-erdos-problem-but-warns-the-win-says-more-about-speed-than-difficulty/
- https://www.wired.com/story/a-new-ai-math-ai-startup-just-cracked-4-previously-unsolved-problems/
- https://www.linkedin.com/pulse/leiden-declaration-hans-martin-will-sqjmf
- https://www.reddit.com/r/mathematics/comments/1tuo7p1/leiden_declaration_on_artificial_intelligence_and/
- https://www.reddit.com/r/LLMmathematics/comments/1uz1ck4/on_the_development_of_ai/
- https://www.nature.com/articles/d41586-026-01881-2.pdf
- https://doi.org/10.1038/s41586-026-10549-w
- https://www.renyi.hu/en/news/community-initiative-mathematicians-already-endorsed-international-mathematical-union
- https://wwwhatsnew.com/2026/06/04/declaracion-leiden-ia-amenaza-matematicas-autonomia-pruebas-2026/
- https://www.dongascience.com/en/news/78235
- https://www.reddit.com/r/antiai/comments/1u04sbg/leiden_declaration_on_artificial_intelligence_and/
- https://www.reddit.com/r/math/comments/1tun9zi/leiden_declaration_on_artificial_intelligence_and/
- https://www.reddit.com/r/accelerate/comments/1uyrbcq/gpt56_sol_pro_oneshotted_all_6_imo_problems_this/
- https://simontechcurator.substack.com/p/the-future-one-week-closer-july-17-2026
- https://www.reddit.com/r/math/comments/1ef0cl3
- https://www.reddit.com/r/ProAI/comments/1vah1pc/using_mostly_my_voice_i_solved_one_of/
- https://openai.com/index/introducing-genebench-pro/
- https://gowers.wordpress.com/2026/07/26/thoughts-about-the-leiden-declaration/
- https://www.reddit.com/r/accelerate/comments/1vah1y9/using_mostly_my_voice_i_solved_one_of/
- https://www.techtimes.com/articles/320669/20260715/gpt-56-disproves-statistics-conjecture-90-minutes-exposing-flaw-130000-citation-method.htm
- https://www.reddit.com/r/BlackboxAI_/comments/1nzu1lw
- https://www.zuhd.news/a/2026-05-10-gowers-chatgpt-5-5-pro-mathematics-failure-modes
- https://arxiv.org/abs/2509.18383
- https://openai.com/index/gpt-5-6/
- https://arxiv.org/abs/2511.16072
- https://di.gg/ai/eyrs8su1?rank=8
- https://arxiv.org/abs/2607.15766
- https://arxiv.org/abs/2607.23386
- https://arxiv.org/abs/2607.11079
- https://arxiv.org/abs/2605.26087
- https://arxiv.org/pdf/2402.07138
- https://arxiv.org/pdf/2402.17879
- https://web3.arxiv.org/pdf/2406.08660
- https://arxiv.org/pdf/2310.02277
- https://arxiv.org/pdf/2402.01869
- https://arxiv.org/pdf/2311.09721
- https://openai.com/index/model-disproves-discrete-geometry-conjecture/
- https://openai.com/research/index/milestone/
- https://simonwillison.net/entries/
- https://simonwillison.net/notes/
- https://openai.com/ro-RO/index/model-disproves-discrete-geometry-conjecture/
- https://simonwillison.net/
- https://openai.com/de-DE/index/model-disproves-discrete-geometry-conjecture/
- https://the-decoder.com/google-deepminds-alphaproof-nexus-solves-decades-old-math-problems-for-a-few-hundred-dollars/
- https://openai.com/pt-BR/index/model-disproves-discrete-geometry-conjecture/
- https://openai.com/ca-ES/index/model-disproves-discrete-geometry-conjecture/
- https://openai.com/fr-FR/index/model-disproves-discrete-geometry-conjecture/
- https://openai.com/index/safety-alignment-long-horizon-models/
- https://cdn.openai.com/pdf/74c24085-19b0-4534-9c90-465b8e29ad73/unit-distance-proof.pdf
- https://cdn.openai.com/pdf/74c24085-19b0-4534-9c90-465b8e29ad73/unit-distance-remarks.pdf
- https://cdn.openai.com/pdf/1625eff6-5ac1-40d8-b1db-5d5cf925de8b/unit-distance-cot.pdf
- https://cdn.openai.com/pdf/00191fd2-3b93-47a3-aff3-6e8bcf787959/unit-distance-proof.pdf
- https://the-decoder.com/openai-launches-o1-and-chatgpt-pro-for-200-per-month/
- https://the-decoder.com/openai-reportedly-plans-to-unveil-ph-d-level-super-agents-at-the-end-of-january/
- https://the-decoder.com/former-go-champion-lee-sedol-still-seems-to-be-struggling-with-ai-defeat/
- https://the-decoder.com/teslas-optimus-robot-this-is-how-experts-see-elon-musks-ai-robot/
- https://dchan.qorigins.org/qresearch/res/6439957.html
- Ten advances in mathematics and theoretical computer science | OpenAI
- Our First Proof submissions | OpenAI














Leave a Reply