← all news

A math benchmark of unsolved Erdős problems that models can't fake

AI · · · source (epoch.ai)

Epoch AI has built a math benchmark that is hard to fake. FrontierMath Erdős is a set of 68 genuinely unsolved problems first posed by Paul Erdős, chosen by the mathematician Thomas Bloom from a pool of 652 open problems and formalized in the Lean proof assistant. Because the answers are checked by Lean rather than by a human grader, a model cannot bluff its way to a passing score: any proof it writes is verified programmatically, so a wrong or hand-wavy argument simply fails. Bloom picked roughly the hardest and most significant tenth of the available problems, and 18 of the 68 were newly formalized for this project on top of 50 that came from Google's Formal Conjectures work.

The scores are sobering. Under the official budget of $300 and 72 hours per problem, only GPT-6 Astra solved anything at all, and it managed two of the 68, about 3%. It disproved one problem for $218 in 15 hours and proved another for $247 in 16 hours. GPT-5.6 Sol, GPT-5.5, Claude Fable 5.1 and Claude Fable 5 all scored zero. When Epoch removed the budget cap and let Astra run further, it reached five solved problems in total but burned more than $220,000 in compute doing it. The full announcement is on Epoch's site.

Why it matters

If you are trying to judge whether AI can do real mathematical research, this gives you a number that resists gaming: today's best model solves about 3% of curated open problems, and only by spending real money. It is a sober baseline to hold against the louder claims about AI cracking major math.

BenchmarksEpoch AIReasoning