
📺 Today’s recommended deep-dive video: https://www.youtube.com/watch?v=wNWz5Hbh5VQ
Beyond the Gold Medal: How OpenAI Reasoning Solved an 80-Year-Old Math Mystery
For decades, the Erdős unit distance conjecture stood as a pillar of combinatorial geometry—until an AI model took it for a “test drive.” This is the inside story of how OpenAI’s reasoning researchers transitioned from high school math competitions to peer-reviewed breakthroughs that stunned the mathematical community.
Core Question: Can large language models equipped with test-time compute transcend human pattern matching to achieve original scientific discovery?
Highlights
- The transition from solving Olympiad problems to disproving the 80-year-old Erdős unit distance conjecture.
- The mechanism of “test-time compute”: how giving a model more time to think improves accuracy exponentially.
- The surprising creative cross-pollination of number theory and geometry performed by the AI during its 125-page “chain of thought.”
- Why the role of the mathematician is shifting from manual calculation to high-level “theory digestion” and AI collaboration.
⏱️ Reading time: approx. 6 minutes · Saves you about 35 minutes vs. watching.
Want to take notes while watching? Click the image below and let AI Notebook capture the key points for you 👇
The Leap from Competitions to Discovery
From IMO Gold to Research Frontiers
For years, the AI community viewed the International Math Olympiad (IMO) as the ultimate “grand challenge” for machine intelligence.
Researchers initially expected models to achieve IMO gold by 2026, but the pace of development shattered those timelines. By June, OpenAI had already secured a gold-medal performance, prompting the reasoning team to look toward harder, unsolved horizons. They turned their attention to the Erdős unit distance conjecture, an 80-year-old problem in combinatorial geometry that asks how many pairs of points in a set can be exactly one unit apart. The conjecture suggested a specific upper limit based on a square grid arrangement, but the team suspected a model might find a better way.
This wasn’t a standard benchmark test; it was an attempt to see if a general-purpose model could actually contribute to the frontier of mathematical research.

💡 Digging Deeper
Q: Why were the researchers so excited about this specific Erdős problem?
A: It was a “major open problem” with a $500 reward attached—a central question in discrete geometry that many human mathematicians had failed to solve for eight decades.
Q: How did the researchers verify that the model’s proof wasn’t just “hallucinated slop”?
A: They shared the result with internal math experts who initially dismissed it as impossible, but after 24 hours of scrutiny, they could find no errors in the logic.
Q: Was the model specialized specifically for mathematics?
A: No, it was a general-purpose model. The researchers describe it as a “test drive” for a reasoning engine that can also code and browse the web.
Scaling the “Thinking” Process
The Power of Test-Time Compute
The secret sauce behind this breakthrough isn’t just a larger dataset, but the implementation of test-time compute. Unlike older models that respond instantly, this reasoning model is given the computational “breathing room” to explore multiple paths, check its own definitions, and iterate on its logic before committing to a final answer.
As compute budget increases, accuracy follows an exponential curve, reaching nearly 50% on problems once considered impossible.
One of the most human-like behaviors observed was the model’s grounding process. Before tackling the complex geometry, the model actually visited the Cambridge Dictionary online to verify the definition of the word “unit.” This step ensured that its mathematical foundation was rock-solid before it attempted to bridge number theory with geometry in a way that surprised even the seasoned researchers on the team. The resulting proof was not a simple calculation but a sophisticated construction using high-powered number theory that human mathematicians have already begun using to knock down other open problems.

💡 Digging Deeper
Q: What did the model do that was considered “creative”?
A: It connected class field theory to combinatorial geometry—a bridge that some humans suspected existed but few had the insight or technical stamina to execute perfectly.
Q: How long was the model’s “internal monologue” for this proof?
A: The chain of thought spanned roughly 125 pages, containing both dead ends and brilliant flashes of insight.
Q: Does this mean AI is better than humans at all math now?
A: Not yet. While it excels at problem-solving, it still struggles to invent entirely new branches of mathematics or build comprehensive new theories from scratch.
The Future of the Human-AI Collaboration
A New Toolkit for Science
While some fear AI will render mathematicians obsolete, the researchers argue the opposite: it acts as a massive productivity multiplier. In the week following the discovery, human mathematicians took the model’s creative construction and applied it to other open problems, successfully disproving a product conjecture for real numbers. This symbiosis proves that AI can handle the “heavy lifting” of exploration while humans focus on high-level theory building and “digesting” the new ideas the AI generates.
The goal isn’t to race through every Erdős problem, but to empower every scientist with top-tier reasoning tools.
Looking forward, the team eyes even bigger prizes like the P vs. NP problem. While currently out of reach, the rapid scaling of reasoning suggests that we are entering an era where AI doesn’t just solve problems we give it—it helps us build the very tools we need to understand the universe. The researchers suggest a “double your trust” approach: trust the model a little more each month, see where it fails, and adjust your workflow to maximize its growing capabilities.
💡 Digging Deeper
Q: What is the “dinosaur” habit the researchers warn against?
A: Not trusting the models enough. Researchers who are used to models from six months ago often under-utilize the current reasoning capabilities by over-decomposing problems.
Q: Can this technology be applied to physics or chemistry?
A: Yes. The team believes this general reasoning ability will accelerate all fields of science, particularly in areas like quantum error correction and material science.
Key Takeaways
The breakthrough disproving the Erdős unit distance conjecture marks a shift in AI from a “retrieval engine” to a “discovery engine.” By leveraging test-time compute, OpenAI models can now bridge disparate mathematical fields—such as number theory and geometry—to solve problems that have remained open for nearly a century. This isn’t just about math; it is a proof of concept for “AI for Science,” where the bottleneck is no longer human calculation speed but our ability to digest and direct AI-generated insights.
The relationship between mathematicians and AI is evolving into a collaborative partnership. AI provides the “construction” and the brute-force reasoning, while humans provide the intuition to apply those constructions to broader theories. As models begin to “reason” for longer durations, the boundary of what is considered “unsolvable” is receding faster than even the researchers anticipated.
Q&A
Q1: Did the model use specialized tools like Lean or Python for the Erdős proof?
A: It functioned as a general ChatGPT-style setup. While it can execute Python code to verify steps, the core breakthrough came from its internal reasoning and browsing capabilities.
Q2: How much did the model “think” compared to a human?
A: The researchers noted that as the compute budget grows, the model’s accuracy on these hard problems grows significantly. It essentially performs a massive parallel search for logical truth.
Q3: What was the most surprising “human” behavior the model showed?
A: Its “chain of thought” included 125 pages of ideas. Even when it pursued paths that didn’t work out, it showed a level of creativity that researchers found “dreamlike.”
Q4: Is the International Math Olympiad (IMO) still a challenge for these models?
A: The researchers consider IMO-level problems to be in the “rearview mirror.” Current models are now tackling peer-reviewed, research-level mathematical conjectures.
Q5: Will AI eventually solve the P vs. NP problem?
A: The team is cautious. Solving P vs. NP likely requires building an entirely new theory, which is currently a bridge too far for models that excel primarily at problem-solving.
Q6: How should a researcher start using these new reasoning models?
A: The team suggests asking the “boldest questions possible” rather than trying to manually decompose problems, as the model’s reasoning may find a better path than the human’s prior assumptions.
Q7: Will this impact cryptography?
A: Potentially. AI might be used to stress-test the foundations of cryptography, either proving that current protocols are mathematically secure or finding loopholes in our conjectures.
