arXiv · 2609.24555
The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
Abstract
We introduce the Endless Exam, a benchmark for measuring mathematical progress from today's models toward artificial superintelligence through fourteen parameterised construction families. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at $1$. The families draw on open mathematical problems for long-term targets and generate new instances at larger parameters, where compact certificates keep large constructions verifiable. Across eight models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though no evaluated system surpasses a published frontier. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Muhan Zhang. 2026-09-21. The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence. https://arxiv.org/abs/2609.24555
Cite the original work for its findings. Save a collection to share your selection of sources.