Citation and evidence

FrontierMath

41 min full readUpdated 29 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI BenchmarksArtificial IntelligenceMathematics

Cite this article

**

FrontierMath
Overview
Full nameFrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
AbbreviationFrontierMath
DescriptionA benchmark of research-level mathematics problems designed to evaluate advanced mathematical reasoning in AI systems
Release date2024-11-08
Latest versionv2 (2026-06-12, 338 corrected problems); Open Problems pilot (2026-01-27)
Benchmark updated2026-06-12 (v2: errors corrected in 42% of original problems)
AuthorsElliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falkman Olsson, Jean-Stanislas Denain, Anson Ho, Emily de Oliveira Santos, Olli Jarviniemi, Matthew Barnett, Robert Sandler, Jaime Sevilla, Qiuyu Ren, Elizabeth Pratt, Lionel Levine, Grant Barkley, Natalie Stewart, Bogdan Grechuk, Tetiana Grechuk, Shreepranav Varma Enugandla
OrganizationEpoch AI
Technical Details
TypeMathematical Reasoning, Research Mathematics
ModalityText, Code
Task formatOpen-ended problem solving with code execution
Number of tasksv1: 350 (Tiers 1-3: 300; Tier 4: 50); v2: 338 (Tiers 1-3: 295; Tier 4: 43) + 14 Open Problems
Total examples338 problems (v2) plus the Open Problems pilot
Evaluation metricAccuracy, Automated verification
DomainsNumber theory, Combinatorics, Group theory, Algebraic geometry, Real analysis, Category theory, Probability theory, Algebraic topology, and 20+ additional fields
LanguagesEnglish
Performance
Human performance~90% (expert mathematicians with days of effort)
Baseline<2% (most models at launch, November 2024)
SOTA score (v1)52.4% (Tiers 1-3), 39.6% (Tier 4)
SOTA model (v1)GPT-5.5 Pro
SOTA date2026-04-23
Tier 4 (v2) leaders97.6% for GPT-6 Astra (OpenAI-reported, 2026-09-03); 87.8% for Claude Fable 5.1 and Claude Fable 5 (highest on Epoch AI's hub as of 2026-09-03)
SaturatedTiers 1-3: no. Tier 4: OpenAI describes it as "saturated" by GPT-6 Astra (company-reported); no independent Tier 4 run of Astra had been published as of 2026-09-03
Resources
WebsiteOfficial website
PaperPaper
DatasetDownload
LicenseProprietary (partial public release)

FrontierMath is an advanced mathematical reasoning benchmark created by Epoch AI in collaboration with over 60 expert mathematicians, including Fields Medalists Terence Tao, Timothy Gowers, and Richard Borcherds. First published on November 8, 2024, FrontierMath consists of hundreds of original, research-level mathematics problems designed to test the outer limits of artificial intelligence systems' mathematical capabilities. At launch, every frontier AI model scored below 2% on the benchmark. By April 2026, the best-performing model, OpenAI's GPT-5.5 Pro, solved 52.4% of Tier 1-3 problems and 39.6% of Tier 4 problems, marking more than a 25-fold improvement in under two years[1][15]. On September 3, 2026, OpenAI reported 97.6% on the corrected Tier 4 (v2) set for GPT-6 Astra, a company-reported figure that Epoch AI had not independently run as of that day; the highest Tier 4 (v2) score on Epoch AI's own hub at the time was 87.8%, for Claude Fable 5.1 and Claude Fable 5[20][22].

The project also includes FrontierMath: Open Problems, a pilot collection of 14 genuinely unsolved mathematical problems whose solutions, if found, would advance the state of human mathematical knowledge[2]. On June 12, 2026, Epoch AI released FrontierMath v2 after an AI-assisted audit found small but critical errors in 42% of the original problems; the v2 release corrected 135 problems and removed 12, leaving a 338-problem set (295 in Tiers 1-3 and 43 in Tier 4)[16][18][19].

What problem does FrontierMath solve?

By 2024, the most widely used mathematical benchmarks for AI had become saturated. Models routinely scored above 95% on GSM8K (grade-school math), above 90% on the MATH dataset (competition-level problems), and 70-90% on AIME-style questions[3]. These high scores made it difficult to distinguish between models or to measure genuine progress in mathematical reasoning.

Epoch AI, a nonprofit research organization focused on tracking AI progress, set out to build a benchmark that would remain challenging for years. The core idea was straightforward: recruit active research mathematicians to write problems drawn from their own fields, problems that require hours or days of expert effort and whose answers can be checked automatically by a computer program.

Elliot Glazer, the project's lead mathematician, holds a Ph.D. in mathematics from Harvard, where he studied set theory under Hugh Woodin. He was joined by Tamay Besiroglu, Epoch AI's associate director, and Ege Erdil as the three core contributors. The broader team eventually grew to include over 60 mathematicians from institutions such as MIT, Harvard, Princeton, Stanford, Cambridge, Oxford, the University of Leicester, King's College London, Cornell, UC Berkeley, and Bristol University, among others. Fourteen IMO gold medalists and three Fields Medal recipients participated in problem creation or review[4].

How is FrontierMath structured?

FrontierMath has expanded since its initial release into four distinct components, each targeting a different level of mathematical difficulty.

Tiers 1-3 (base set)

The original base set contains 300 problems spanning difficulty from advanced undergraduate to early postdoctoral level. This set forms the core benchmark used in most published evaluations. Problems are classified using the Mathematics Subject Classification (MSC2020) system and cover virtually every major branch of modern mathematics[4]. After the June 2026 v2 correction, the base set comprises 295 problems (123 corrected and 5 removed)[18][19].

Tier 4 (expansion set)

Released on July 1, 2025, Tier 4 adds 50 exceptionally difficult research-level problems to the benchmark. These problems were largely designed or refined during a symposium attended by leading mathematicians, where each problem was tested and approved by a panel of experts. Of the 50 Tier 4 problems, 2 are public and 48 are private. Even the strongest AI systems as of mid-2025, including OpenAI's o4-mini, Anthropic's Claude Opus 4, and Google's Gemini 2.5 Pro, achieved only single-digit success rates on Tier 4[5]. By April 2026, GPT-5.5 Pro had pushed the Tier 4 frontier to 39.6%[15]. The v2 update reduced Tier 4 to 43 problems (12 corrected and 7 removed)[18][19]. Two of the 43 are public, and Epoch AI's hub scores models on the remaining 41 private problems[22]. On September 3, 2026, OpenAI reported 97.6% on the v2 Tier 4 set for GPT-6 Astra (see below)[20].

Open Problems

On January 27, 2026, Epoch AI launched FrontierMath: Open Problems, a pilot benchmark of 14 genuinely unsolved mathematical research problems. Unlike the main benchmark, where problems have known solutions that an expert created, these are problems that professional mathematicians have attempted and failed to solve. The pilot set tilts toward combinatorics and number theory, where the most problems amenable to automatic verification were found[2].

Each open problem includes a difficulty estimate from its contributor. Estimated solving times range from one to four weeks at the low end to three to ten years at the high end. The number of serious human attempts per problem ranges from two or three mathematicians to over fifty. Significance ratings span from "moderately interesting results" to "major breakthroughs"[2].

Two problems added to the benchmark in February 2026 illustrate the scope: finding a Hadamard matrix of order 668 (the smallest order for which none is known) and proving that certain "small" Diophantine equations have infinitely many solutions[6].

FrontierMath Erdős

On September 1, 2026, Epoch AI announced FrontierMath Erdős, a set of 68 problems posed or studied by Paul Erdős, all open as of August 2026, selected by Thomas Bloom (who maintains erdosproblems.com) as both significant and difficult. Unlike the other components, the problems are formalized in Lean, and a problem counts as solved only when a model produces a Lean proof or disproof that passes verification; 50 of the 68 formalizations come from Google's Formal Conjectures project and the remaining 18 were AI-formalized and are still under expert review. Epoch AI runs each model once per problem with a budget of $300 and 72 hours per attempt, gives the models an offline collection of mathematics papers and a computer algebra system rather than internet access, and open-sourced the harness. Epoch AI lists three caveats: formalization is a separate burden bolted onto the mathematics, published solutions will eventually contaminate training data, and Erdős problems are not a sample of all mathematics[24].

How are FrontierMath problems designed?

Core requirements

Every FrontierMath problem must satisfy four requirements before it enters the benchmark[4]:

RequirementDescriptionPurpose
OriginalityProblems build on existing ideas in novel, non-obvious ways through clever adaptations or innovative combinationsPrevents data contamination from training sets
Automated verifiabilitySolutions must be computable and expressible as Python objects or SymPy structures (integers, symbolic expressions, matrices, sets)Allows scalable, objective evaluation
GuessproofnessLess than 1% probability of arriving at the correct answer without performing the required mathematical workEnsures models cannot succeed through random guessing or superficial heuristics
Computational tractabilitySolution verification scripts must run in under one minute on standard hardwareKeeps evaluation practical

Expanded article table

Difficulty rating system

Each problem is rated along three dimensions by its creator and at least one peer reviewer[4]:

DimensionScaleDescription
Background knowledge1-51 = high school level; 2 = early undergraduate; 3 = late undergraduate; 4 = graduate; 5 = research level
CreativityHours (unbounded)Time an expert in the relevant field would need to identify the key solution ideas
ExecutionHours (unbounded)Time to compute the final answer once the key ideas are identified, including writing any necessary code

Expanded article table

The authors note that these ratings provide rough guidance rather than definitive claims, since problems can become easier once a specific technique is known, and multiple solution paths of varying difficulty may exist[4].

Mathematical domain coverage

The benchmark spans most major branches of modern mathematics. The distribution of problems by MSC2020 primary classification is as follows[4]:

MSC CodeFieldShare of problemsInvolvement in multi-domain problems
11Number theory17.8%44% of all problems involve number theory
05Combinatorics15.8%39% of all problems involve combinatorics
20Group theory8.9%22% of all problems involve group theory
60Probability theory5.1%-
15Linear algebra4.8%-
14Algebraic geometry4.8%-
33Special functions4.8%-
55Algebraic topology3.1%-
12Field theory2.4%-
30Complex analysis2.4%-
68Computer science2.4%-
18Category theory2.4%-
57Manifolds and cell complexes2.1%-
13Commutative algebra2.1%-
Other17 additional fields21.1%Includes PDEs, differential geometry, harmonic analysis, statistical mechanics, and more

Expanded article table

Notably, 13% of problems combine number theory and combinatorics, 9% combine combinatorics and group theory, and 8% combine number theory and group theory. Over 200 distinct solution techniques are represented across the benchmark, and even the most common techniques (generating functions, recurrences, special functions) each appear in fewer than 5% of problems[4].

How are problems created and vetted?

Creation pipeline

The process for creating and reviewing FrontierMath problems involves multiple stages[4]:

StageProcessQuality control
Problem designExpert mathematicians create original problems in their research areasMust satisfy all four core requirements
Solution developmentAuthors write a solution script in Python that computes the answerScript must terminate in under one minute
Verification designDevelop automated checking methods using exact matching, SymPy evaluation, or computational verificationEnsure answers are unambiguous
Blind peer reviewAt least one domain expert mathematician reviews each problem without knowledge of the solution approachReviewers assess correctness, ambiguity, guessproofness, and difficulty ratings
Second-round reviewA random subset of 25 problems receives an additional blind reviewProvides error rate estimates
Error correctionProblems flagged during review are revised or removedEstimated error rate: roughly 10% (1 incorrect answer found in 25 reviewed problems)
Final validationComplete verification testing on all accepted problemsConfirms automated checking works reliably

Expanded article table

Anti-contamination measures

Because the value of the benchmark depends on problems being unknown to AI training pipelines, Epoch AI employs several security measures[4]:

  • All problems are original and previously unpublished
  • Communication with contributors uses encrypted channels
  • Problem files are shared via password-protected archives
  • A core mathematician team manually checks problems against mathematics websites, repositories, and academic publications
  • Plagiarism detection tools (Quetext and Copyscape) scan the full dataset
  • The majority of the benchmark remains private, with only a handful of sample problems released publicly

Guessproof verification

Each problem undergoes a guessproofness check to confirm that the answer space is large enough (typically exceeding 10^6 possibilities) and that no obvious patterns would allow a model to stumble on the correct answer. Problems typically require large, non-obvious numerical answers or complex mathematical objects as solutions. The target is a less than 1% success rate for random or heuristic guessing[4].

How are models evaluated?

Interactive environment

Models are evaluated in an interactive Python environment. The evaluation framework gives each model access to the following capabilities[4]:

CapabilityDescription
Code executionWrite and run Python code to perform calculations
Library accessUse standard mathematical libraries (SymPy, NumPy, SciPy, etc.)
Iterative problem solvingMultiple attempts are allowed within the token budget
Result verificationModels can check intermediate results before final submission

Expanded article table

For Tier 4 evaluations, models receive a 1,000,000-token hard limit with a 660,000-token warning threshold. The model submits a Python function that returns its answer after reasoning and code execution[5].

Answer verification

When a model submits its answer, verification proceeds automatically[4]:

MethodDescriptionExample
Exact integer matchingCompare submitted integer to known answer"The answer is 3677073"
SymPy symbolic evaluationCheck if the difference between submitted and known expressions simplifies to zeroPolynomial equality
Computational object verificationVerify properties of submitted mathematical structuresCheck that a submitted matrix satisfies required group properties
Numerical toleranceFor floating-point answers, check within a specified toleranceApproximation results

Expanded article table

The model's code must include a specific marker comment (# This is the final answer), save the result using Python's pickle module, and be fully self-contained[4].

How have AI models performed on FrontierMath?

Timeline of AI performance on FrontierMath (Tiers 1-3)

The following table shows how model performance has evolved since the benchmark's release. These figures are scored against the v1 dataset; v2 scores released after June 12, 2026 are not directly comparable[1][7][8][9][10][15]:

ModelOrganizationScoreDateNotes
GPT-5.5 ProOpenAI52.4%April 2026v1 SOTA; 39.6% on Tier 4
GPT-5.5OpenAI51.7%April 2026Released April 23, 2026
GPT-5.4 ProOpenAI~50%March 2026Previous SOTA; 38% on Tier 4
Claude Opus 4.7Anthropic~44%April 2026Released April 16, 2026; adaptive thinking variant
GPT-5.2 (Thinking)OpenAI40.3%Late 2025First model above 40%
GPT-5.1OpenAI26.7%2025Multiple variants at same score
GPT-5OpenAI26.3%2025-
GPT-5 miniOpenAI22.1%2025-
o3 (public release)OpenAI~10% (Epoch AI), 25.2% (OpenAI internal)April 2025 / December 2024Score discrepancy became controversial (see below)
Grok 4xAI~14%2025-
Gemini 2.5 ProGoogle DeepMind~11%2025-
o3-miniOpenAI8.9-9.2%2025Medium reasoning setting
Claude Opus 4.1Anthropic~7%2025Epoch AI evaluation
o1OpenAI5.5%2025-
DeepSeek R1DeepSeek5.2%2025Open-source leader at the time
Gemini 2.0 Flash ThinkingGoogle2.6%2025Experimental version
Claude 3.5 SonnetAnthropic<2%November 2024Initial evaluation
GPT-4oOpenAI<2%November 2024Initial evaluation
o1-previewOpenAI<2%November 2024Initial evaluation
Gemini 1.5 ProGoogle<2%November 2024Initial evaluation
Grok 2 BetaxAI<2%November 2024Initial evaluation

Expanded article table

The top three Tier 1-3 models (GPT-5.5 Pro, GPT-5.5, GPT-5.4 Pro) cluster within 2.4 percentage points, prompting commentators to describe the benchmark as approaching saturation among frontier models even as more than 45% of v1 problems remain unsolved[9].

Tier 4 performance

Tier 4 scores are reported separately due to the significantly higher difficulty[5][15]:

ModelScoreNotes
GPT-5.5 Pro39.6%April 2026; v1 Tier 4 SOTA
GPT-5.4 Pro~38%March 2026
GPT-5.535.4%April 2026; OpenAI reported
Claude Opus 4.722.9%April 2026; per OpenAI's GPT-5.5 comparison
Gemini 3 Pro19% (+/- 6%)3 of 48 samples failed due to API errors
Grok 42% (+/- 2%)8 of 48 samples had API errors
DeepSeek V3.2 (Thinking)~2%Only Chinese-origin model to score above zero on Tier 4

Expanded article table

Tier 4 (v2) results on Epoch AI's hub

After the June 12, 2026 correction, Epoch AI re-ran models on the 41 private problems of the corrected Tier 4 set and publishes the results, with standard errors, on its AI Benchmarking Hub. Scores on the corrected set run far higher than the v1 figures above and are not comparable with them. The leading entries as of September 3, 2026 were[22]:

Model (setting)Score (Epoch AI, v2 private set)Run started
Claude Fable 5.1 (max)87.8% (+/- 5.2)September 1, 2026
Claude Fable 5 (max)87.8% (+/- 5.2)June 9, 2026
GPT-5.6 Sol (max)82.9% (+/- 5.9)July 9, 2026
GPT-5.6 Sol (pro, max)80.5% (+/- 6.3)July 9, 2026
GPT-5.5 Pro (xhigh)78.0% (+/- 6.5)June 12, 2026
AI co-mathematician (Google DeepMind)75.6% (+/- 6.7)June 12, 2026
Claude Opus 5 (max)73.2% (+/- 7.0)July 24, 2026
GPT-5.5 (xhigh)72.5% (+/- 7.1)June 11, 2026
GPT-5.6 Terra (max)70.7% (+/- 7.2)July 9, 2026
GPT-5.6 Luna (max)61.0% (+/- 7.7)July 9, 2026
GPT-5.4 Pro (xhigh)58.5% (+/- 7.8)June 13, 2026

Expanded article table

The hub's download listed 54 Tier 4 (v2) runs on that date. It contained no Tier 4 run of GPT-6 Astra[22].

GPT-6 Astra (September 2026)

On September 3, 2026, OpenAI launched GPT-6 Astra and reported a FrontierMath result in two forms. The introduction of the launch post says Astra "saturates FrontierMath Tier 4 with a 98% score", while the post's Academic results table lists "FrontierMath Tier 4 (v2)" at 97.6% for Astra against 83.0% for GPT-5.6 Sol, 87.8% for Claude Fable 5.1, 87.8% for Claude Fable 5, and 73.2% for Claude Opus 5, with no Gemini 3.8 Flash entry. The table's note says the scores are "the maximum at any effort" and that GPT evaluations were run in OpenAI's research environment or via its API[20]. A pre-briefing table published by The New Stack shows the same FrontierMath row, and VentureBeat and The New Stack both reported the 97.6% figure[21][25]. The "(v2)" label refers to the corrected 43-problem Tier 4 set released on June 12, 2026 (see below); the post does not say which problems were used, how many attempts were run, or whether the two public problems were included[20]. All of these numbers are OpenAI's own; the launch post does not link to an Epoch AI evaluation.

ModelOpenAI's launch table (Tier 4 v2)Epoch AI hub, September 3, 2026
GPT-6 Astra97.6% (98% in the post's introduction)Not listed
Claude Fable 5.187.8%87.8% (max)
Claude Fable 587.8%87.8% (max)
GPT-5.6 Sol83.0%82.9% (max); 80.5% (pro, max)
Claude Opus 573.2%73.2% (max)

Expanded article table

The New Stack's Frederic Lardinois wrote that the 97.6% result "appears to cover the 41 private problems in the 43-problem tier" and reminded readers that Epoch AI, "which runs the benchmark, says OpenAI funded its development and has exclusive access to part of it"[21]. Both points check out against Epoch AI's own pages. Epoch AI's hub states that two Tier 4 problems are public and that, unless stated otherwise, its numbers cover the private set; on 41 problems, 97.6% is the rounded value of 40 solved out of 41, and the other four scores in OpenAI's row are the same values Epoch AI's hub lists for those models, apart from GPT-5.6 Sol (OpenAI 83.0%, Epoch AI 82.9%), so the row is consistent with the 41-problem private set even though OpenAI does not say so[20][22]. On funding and access, Epoch AI's conflict-of-interest statement says OpenAI commissioned the 300 core problems and the 50 Tier 4 problems, has access to all core problem statements and solutions except 53 solutions withheld in the February 2025 version, and has access to 30 of the 50 original Tier 4 problems, with the remaining 20 reserved as a holdout set; Epoch AI retains the right to evaluate any model on the full dataset[23][26]. Epoch AI has not published how the holdout split maps onto the corrected 43-problem v2 set.

As of September 3, 2026, Epoch AI had not published an independent Tier 4 evaluation of GPT-6 Astra: its Benchmarking Hub listed no Astra run on any FrontierMath tier, and its publications feed carried nothing on the launch[22]. The only Epoch AI figure for the model is on FrontierMath Erdős, where, in the September 1 announcement co-written by Tom Adamczewski and Epoch AI's head of benchmarks Greg Burnham, a pre-release version of GPT-6 Astra solved 2 of the 68 problems (3%) in the benchmark run: it disproved problem 74 by finding a counterexample ($218, 15 hours) and proved problem 126 ($247, 16 hours), while GPT-5.6 Sol, GPT-5.5, Claude Fable 5.1 and Claude Fable 5 each scored 0%. In further, non-benchmark attempts with larger budgets and varied scaffolds, the same pre-release Astra solved three more (problems 1, 548 and 571), five of 68 in total, at more than $220,000 of compute against roughly $20,000 for the benchmark run itself; Epoch AI stressed that these extra attempts are not a FrontierMath Erdős score[24]. OpenAI's launch post carries a pull quote attributed to Burnham, "The story is: end of one era, start of another", placed in its mathematics section without further context[20].

The same launch post frames the Tier 4 result alongside Astra's research mathematics: the ten results OpenAI attributed to an internal version of Astra on August 1, 2026, and two new prime-gap manuscripts released with the launch, "Improved short gaps between primes" (infinitely many pairs of consecutive primes within 186 of each other) and "Improved long gaps between primes", each stating that "the proof is due to GPT 6 Astra"[20][27][28][29]. Those results, and the debate over how much of them was the model's own work, are covered on the GPT-6 Astra page at /wiki/gpt_6_astra.

Initial model behavior patterns

In the original November 2024 evaluation, the paper's authors documented several behavioral patterns across the six tested models[4]:

  • o1-preview averaged 1.29 responses per problem, while Grok 2 Beta averaged 3.81 responses per problem
  • o1-preview and Gemini 1.5 Pro tended to submit answers before seeing experimental results, even when the evaluation framework encouraged iterative testing
  • Claude 3.5 Sonnet, GPT-4o, and Grok 2 Beta exceeded the 10,000-token limit in over 45% of attempts
  • Gemini 1.5 Pro hit the token limit in only 16.8% of attempts, using roughly 6,000 tokens on average compared to 12,000-17,000 for other models
  • Across five runs per model per problem, only four problems total were solved by at least one model; o1-preview was the only model to solve any problem on all five runs

What do expert mathematicians say about FrontierMath?

Four prominent mathematicians were interviewed for the FrontierMath paper: Terence Tao (2006 Fields Medalist), Timothy Gowers (1998 Fields Medalist), Richard Borcherds (1998 Fields Medalist), and Evan Chen (IMO coach and benchmark co-author). Their comments offer a window into how professional mathematicians view the benchmark's difficulty and significance[4].

Terence Tao

Tao contributed several problems to the benchmark and reviewed others. He described the problems as "extremely challenging" and predicted the benchmark would "resist AIs for several years at least." On the scarcity of relevant training data, Tao observed that for many FrontierMath problems, the relevant material is "almost nonexistent... you're talking like a dozen papers with relevant things"[4].

Tao suggested that human experts working alongside AI systems could tackle FrontierMath problems within about three years, noting that guiding current AI to correct solutions takes "about five times as much effort" as solving the problems directly. He expected this ratio to improve and eventually drop below 1 for certain problems within a few years, given sufficient tooling and capability improvements[4].

On practical considerations, Tao remarked that if AI tools require "three days of compute off of all of Google to solve each problem... that's less of a useful tool"[4].

Timothy Gowers

Gowers reported that "all of the problems I looked at were not really in my area and all looked like things I had no idea how to solve." He emphasized that the problems "appear to be at a different level of difficulty from IMO problems," requiring familiarity with "the tricks of the trade of some particular branch of maths," a kind of domain knowledge that is hard to acquire without substantial, specialized training data[4].

Gowers also offered a practical vision for AI in mathematics, suggesting that AI systems could help with "slightly boring bits of doing research where you, for example, make some conjecture that would be useful, but you're not quite sure if it's true... it could be a very, very nice time saving device"[4].

Richard Borcherds

Borcherds was described in the paper as "the most bullish" among the interviewees about AI's potential in mathematics. He did note, however, that the benchmark problems "aren't quite the same as coming up with original proofs," drawing a distinction between solving a problem with a known answer and generating new mathematical knowledge[4].

Evan Chen

Evan Chen, a well-known mathematics educator and IMO coach who also co-authored the FrontierMath paper, published a separate blog post analyzing the benchmark's design philosophy. He noted that FrontierMath inverts two of the three desirable properties of traditional competition problems (like those at the IMO or Putnam exam). While FrontierMath retains the requirement for creative insight, it deliberately abandons the simplicity requirement and assumes the solver has "access to a Python console and a lot of reference text." Chen praised the authors for being "pretty ruthless about rejecting problems for which they felt it was possible to guess the answer" through engineer's induction[11].

Chen identified a key advantage of FrontierMath's design: its ability to use "easily verifiable solutions" through code implementation, similar to the International Olympiad in Informatics or Project Euler. This contrasts with pencil-and-paper competitions where human coordinators must evaluate proofs[11].

What was the o3 score controversy?

OpenAI's initial claim

On December 20, 2024, OpenAI announced its o3 reasoning model and reported a 25.2% score on FrontierMath, a dramatic leap from the previous best of under 2%. This result was highlighted during the o3 launch event as evidence of a breakthrough in mathematical reasoning[7].

Epoch AI's independent evaluation

On April 18, 2025, Epoch AI published its own independent evaluation of the publicly released o3 model, reporting a score of approximately 10%, significantly below OpenAI's claim. Epoch AI identified several factors that could explain the discrepancy[8]:

FactorOpenAI's testing (December 2024)Epoch AI's testing (April 2025)
Model versionPre-release internal versionPublic release version, "tuned for chat/product use"
Compute resources"Aggressive test-time compute"Standard compute tiers
Problem set180 problems (frontiermath-2024-11-26)290 problems (frontiermath-2025-02-28)
ScaffoldingInternal advanced scaffoldPublic API scaffold

Expanded article table

Epoch AI noted: "The difference between our results and OpenAI's might be due to OpenAI evaluating with a more powerful internal scaffold, using more test-time computing, or because those results were run on a different subset of FrontierMath"[8].

Funding disclosure controversy

The o3 announcement also triggered scrutiny of the financial relationship between OpenAI and Epoch AI. On the same day o3 was announced (December 20, 2024), Epoch AI disclosed that OpenAI had funded the creation of FrontierMath. Several problems quickly emerged[12][13]:

  • OpenAI had visibility into many of the problems and solutions in the benchmark before the public announcement
  • The more than 60 contributing mathematicians were not informed of OpenAI's involvement or exclusive early access
  • Six mathematicians who contributed significantly to the benchmark confirmed to a Stanford PhD student that they were unaware OpenAI would have exclusive access
  • Epoch AI's associate director acknowledged being "restricted from disclosing the partnership until around the time o3 launched" and stated that "in hindsight we should have negotiated harder for the ability to be transparent to the benchmark contributors as soon as possible"
  • OpenAI and Epoch AI had a "verbal agreement" that OpenAI would not use FrontierMath's problem set to train its AI models

The controversy drew criticism from multiple outlets. Fortune described it as "manipulative and disgraceful." TechCrunch reported that the benchmarking organization was "criticized for waiting to disclose funding from OpenAI." The incident raised broader questions about independence in AI benchmarking and the risks of conflicts of interest when AI companies fund the benchmarks used to evaluate their own models[12][13].

Epoch AI is primarily funded by Open Philanthropy, and the OpenAI funding for FrontierMath was a separate, project-specific arrangement[12].

Why did Epoch AI correct 42% of FrontierMath problems?

On May 11, 2026, Epoch AI announced that it was "conducting an AI-assisted review of FrontierMath: Tiers 1-4" and that the review had "flagged fatal errors in about a third of problems," adding "we believe most are valid flags." The organization said it would release updated scores on a corrected dataset once a thorough human review was complete[16][17].

The disclosure significantly raised the estimated error rate of the benchmark. The paper's original second-round review of 25 problems had flagged roughly 1 in 25 problems (about 4%) as incorrect; the new AI-assisted pass flagged closer to one-third, an order-of-magnitude increase[4][16]. OpenAI researcher Noam Brown publicly credited GPT-5.5 with producing the first flags, an inversion of the usual relationship in which the benchmark evaluates the model rather than the model evaluating the benchmark[16][17].

FrontierMath v2 (June 12, 2026)

Epoch AI completed the human review and released FrontierMath v2 on June 12, 2026. The final audit found small but critical errors in 42% of the original problems, a far higher rate than the roughly 5% (about 1 in 20) suggested by the earlier human quality reviews[18][19]. According to Epoch AI's account, the project began in April 2026 when OpenAI shared that it had found more errors than expected during an internal review; Epoch AI then ran an independent audit, using frontier models such as GPT-5.5 and Claude Opus 4.7 to flag candidate errors before engaging mathematicians to adjudicate the flags. Almost all of the flagged issues were determined to be real and severe enough to render the affected problems unsolvable as stated[18][19].

The v2 release corrected 135 problems and removed 12, leaving 338 problems in total[18][19]:

Setv1 countCorrectedRemovedv2 count
Tiers 1-33001235295
Tier 45012743
Total35013512338

Expanded article table

Epoch AI framed the correction as routine benchmark maintenance rather than a benchmark failure, characterizing the original error rate as "comparable to error rates in other major ML benchmarks like ImageNet"[18]. The organization cautioned that scores measured on v1 are not directly comparable to scores measured on the corrected v2 set, so leaderboard positions established before June 12, 2026 should be treated with care until models are re-run on v2[18][19].

The review carried several implications:

  • Previously reported scores on Tiers 1-3 and Tier 4 may shift once errored problems are removed or corrected, and Epoch AI cautioned against treating pre-v2 leaderboard positions as final[16][18]
  • A frontier model contributing materially to the audit of its own evaluation set complicates the long-standing principle that benchmarks should be created and graded independently from the systems they measure
  • Commentators noted that GPT-5.5's audit performance was itself a capability demonstration: identifying valid mathematical errors at scale in research-level problems requires the same skills the benchmark is designed to test[17]

Epoch AI maintained that final corrections were made by human mathematicians, not by AI alone, and emphasized that the models were used as filters to surface candidate errors rather than as the arbiters of validity[16][18].

FrontierMath: Open Problems and the first AI solution

The Ramsey hypergraph breakthrough

On March 24, 2026, Epoch AI confirmed that GPT-5.4 Pro had produced a verified solution to a genuinely open mathematical problem on FrontierMath: a Ramsey-style problem on hypergraphs that had remained unsolved since it was posed by mathematicians Will Brian and Paul Larson in a 2019 paper. The solution was first elicited by researchers Kevin Barreto and Liam Price using GPT-5.4 Pro. Problem contributor Will Brian confirmed the solution's correctness, and a write-up is being prepared for publication[14].

This marked the first time an AI model produced a novel solution to an open problem on the FrontierMath benchmark. After the initial solve, several other frontier models also solved the same problem: Claude Opus 4.6 (max), Gemini 3.1 Pro, and GPT-5.4 (xhigh). The fact that multiple models could solve it suggests the problem sat at the boundary of current frontier model capabilities[14].

Broader context

The Ramsey hypergraph result is part of a wider trend. Since Christmas 2025, 15 open mathematical problems have moved from unsolved to solved, with 11 of them (73%) credited to AI involvement. However, Epoch AI also noted that when GPT-5.4 Pro was evaluated on the full set of FrontierMath Open Problems, it "did not solve any problems" other than the Ramsey one, and its novel observations on one other problem were "of a form that the author had anticipated and characterized as relatively uninteresting"[14].

Sample problems

While most problems remain private to prevent contamination, the original paper includes five public sample problems at varying difficulty levels[4]:

ProblemAuthorDifficultyField (MSC)Key techniquesCreativity (hours)Execution (hours)Answer
Testing Artin's primitive root conjectureO. JarviniemiResearch levelNumber theory (11)Frobenius elements, Artin symbols4153,677,073
Find degree 19 polynomialA. KiteResearch levelAlgebraic geometry (14), Group theory (20), Number theory (11)Monodromy, branch loci341,876,572,071,974,094,803,391,179
Prime field continuous extensionsD. ChicharroGraduate levelNumber theory (11)p-adic analysis, recurrences339,811
Coxeter group problemP. EnugandlaGraduate levelGroup theory (20)Coxeter groups, characters23(not disclosed in sample)
Algebraic geometry/number theory problemA. GunningUndergraduate levelAlgebraic geometry (14), Number theory (11)Hasse-Weil bound22(not disclosed in sample)

Expanded article table

These samples illustrate several features of the benchmark: answers are large, non-obvious integers (making them guessproof); problems span multiple mathematical fields; and even the "easiest" problem requires two hours of creative work from an expert.

How does FrontierMath compare with other benchmarks?

Difficulty scaling

BenchmarkAI performance (approximate best)Typical problem levelTypical solving time (human)Primary limitation
GSM8K>95%Grade schoolMinutesSaturated since 2024
MATH>90%High school/competition30 minutesSaturated; data contamination risk
AIME70-90%Competition mathematicsHoursApproaching saturation
MMLU (math subset)>85%Mixed undergraduateVariesNot math-specific
FrontierMath (Tiers 1-3)52.4% (v1)Undergraduate to postdocHours to daysStill challenging; majority unsolved
FrontierMath (Tier 4)39.6% (v1); 87.8% (v2, Epoch AI hub); 97.6% (v2, OpenAI-reported for GPT-6 Astra)Research levelDays to weeksFrontier models pass 70% on v2 while most others stay below 50%; Astra's figure has no independent run[20][22]
FrontierMath (Open Problems)1 problem solvedUnsolved researchWeeks to yearsVirtually all problems remain unsolved

Expanded article table

What sets FrontierMath apart

FeatureFrontierMathTypical math benchmarks
Problem sourceOriginal, unpublished, created by active researchersOften drawn from textbooks, competitions, or publicly available problem sets
Answer verificationFully automated via Python/SymPyOften requires human grading or proof checking
Data contamination riskMinimal (private problem set, encrypted distribution)High (problems publicly available, may appear in training data)
Difficulty rangeUndergraduate through active researchTypically grade school through undergraduate
Time investment per problemHours to days for expertsMinutes to hours
Multi-domain integration44% of problems involve multiple mathematical fieldsMost problems stay within a single topic

Expanded article table

Notable contributors

Fields Medalists

NameFields Medal yearRole
Terence Tao2006Problem creation, review, and interview
Timothy Gowers1998Problem review and interview
Richard Borcherds1998Problem review and interview

Expanded article table

Key team members

NameRoleBackground
Elliot GlazerLead mathematicianPh.D. in mathematics from Harvard (set theory under Hugh Woodin)
Tamay BesirogluAssociate director, Epoch AIPreviously at MIT Future Tech Lab; led strategy for Metaculus
Ege ErdilCore contributorEpoch AI researcher
Evan ChenCo-author and contributorIMO coach, mathematics educator

Expanded article table

Institutional participation

Over 60 mathematicians from leading institutions contributed, including researchers from MIT, Harvard, Princeton, Stanford, Cambridge, Oxford, Cornell, UC Berkeley, King's College London, the University of Leicester, the University of Siegen, ICMC USP (Brazil), and Bristol University, among others.

Implementation details

Evaluation setup

Models interact with a Python environment where they can write and execute code, test hypotheses, and submit answers. A simplified conceptual overview of the evaluation framework:

"""Conceptual evaluation framework (simplified)."""
class FrontierMathEvaluator:
    def evaluate_model(self, model, problem):
        environment = PythonEnvironment()
        max_attempts = 10
        for attempt in range(max_attempts):
            code = model.generate_code(problem, environment.state)
            result = environment.execute(code)
            if model.verify_answer(result, problem):
                return self.check_solution(result, problem.answer)
        return False

Access tiers

Access levelDescriptionHow to obtain
Public samplesSmall set of example problems with full solutionsFree access via epoch.ai/frontiermath
Open Problems verifiersSolution verifiers for the 14 open problemsPartnership with Epoch AI (math@epoch.ai); uniform access fee
Research evaluationFull benchmark evaluation on the private setContact math_evals@epoch.ai
Commercial evaluationModel testing servicePartnership with Epoch AI
Problem contributionSubmit new problems for inclusionExpert mathematician credentials required

Expanded article table

Funding and development

Funding sources

FrontierMath's development has been supported by:

  • Open Philanthropy (primary funder of Epoch AI as an organization)
  • OpenAI (project-specific funding, disclosed December 2024)[12]; Epoch AI's January 2025 clarification first disclosed the arrangement (a 50-question holdout, with OpenAI to receive only the problem statements of the then-forthcoming Tier 4 set), and its standing conflict-of-interest statement now says OpenAI commissioned the 300 core problems and the 50 Tier 4 problems, owns them, and can see all statements and solutions except a holdout (53 core solutions and 20 of the 50 original Tier 4 problems), while Epoch AI keeps the right to evaluate any model on the full set[23][26]
  • Additional academic and industry partners

Ongoing development

InitiativeDescriptionStatus
Problem expansionAdding new problems to Tiers 1-4Ongoing; quarterly updates
Domain coverageExpanding to additional mathematical fields2025-2026
Tier 4 updatesBug fixes and grader corrections (version bumped to 1.1.4 in 2026)Ongoing
AI-assisted error reviewReviewing flagged problems and republishing corrected scoresCompleted; FrontierMath v2 released June 12, 2026 (42% of problems had errors)[16][18]
Open Problems growthExpanding beyond the 14-problem pilot setPlanning stage
FrontierMath Erdős68 open Erdős problems formalized in Lean, solved only by verified Lean proofsLaunched September 1, 2026; five models run, a pre-release GPT-6 Astra at 3%[24]
Verification improvementsRefining automated checking methodsContinuous

Expanded article table

Impact and significance

Influence on AI research

FrontierMath has had a measurable effect on the AI research community since its release:

  • It demonstrated that benchmark saturation on easier datasets (GSM8K, MATH) did not indicate genuine mathematical reasoning capability
  • The benchmark's design principles, particularly its emphasis on originality, automated verification, and guessproofness, have influenced the design of subsequent benchmarks
  • The o3 score controversy prompted broader discussion about transparency in AI benchmarking and the risks of vendor-funded evaluation
  • The Open Problems component established a new category of AI evaluation: testing whether models can advance the frontier of human knowledge, not merely match it
  • The May-June 2026 AI-assisted review and v2 correction illustrated a new dynamic: frontier models becoming capable enough to audit the benchmarks designed to evaluate them[16][17][18]

Progress tracking

The trajectory from under 2% (November 2024) to 52.4% (April 2026) on Tiers 1-3 is one of the fastest rates of improvement on any major AI benchmark. On the v1 set the leading Tier 4 score was below 40% and most models scored in single digits; on the corrected v2 set, Epoch AI's hub listed Claude Fable 5.1 and Claude Fable 5 at 87.8% as of September 3, 2026, and OpenAI reported 97.6% for GPT-6 Astra the same day, a company-reported figure with no independent run[20][22]. Virtually all Open Problems remain unsolved. The June 2026 v2 correction, which revised or removed 42% of the original problems, also means pre-v2 figures should be read as approximate until models are re-scored on the corrected 338-problem set[1][5][14][15][16][18].

Limitations and criticisms

Known limitations

LimitationDescriptionMitigation
Limited public accessMost problems are private to preserve benchmark integrityNecessary trade-off; sample problems are publicly available
Narrow scopeOnly tests mathematical problem-solving; does not assess proof writing, mathematical intuition, or pedagogical abilityComplements other benchmarks
English onlyAll problems are written in EnglishFuture multilingual expansion is planned
Computational biasProblems must have automatically verifiable answers, excluding proof-based and open-ended mathematical reasoningAcknowledged limitation of the automated verification approach
Estimated error rateThe original paper estimated roughly 10% errors; the June 2026 v2 audit corrected or removed 42% of problemsv2 dataset released June 12, 2026 after human verification[4][16][18]

Expanded article table

Criticisms

Several criticisms have been raised since the benchmark's launch:

  1. The funding transparency failure undermined trust in the benchmark's independence, even though the problems themselves were created by independent mathematicians[12][13]
  2. The discrepancy between OpenAI's reported 25.2% and Epoch AI's measured 10% for o3 highlighted the difficulty of comparing results when testing conditions differ[8]
  3. Limited access to the full problem set makes independent replication difficult, though this restriction exists to prevent data contamination
  4. Some mathematicians have questioned whether problems with automatically verifiable answers represent the full range of mathematical reasoning, since much of research mathematics involves constructing proofs rather than computing specific values[4]
  5. The June 2026 disclosure that 42% of the original Tier 1-4 problems contained errors raised fresh questions about how confidently pre-v2 year-over-year scores can be compared[16][17][18]

See also

References

  1. ^1 ^2 ^3Epoch AI. "FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI." epoch.ai/frontiermath
  2. ^1 ^2 ^3Epoch AI. "Introducing FrontierMath: Open Problems." Epoch AI Substack, January 2026. epochai.substack.com/...frontiermath-open-problems
  3. ^VentureBeat. "AI's math problem: FrontierMath benchmark shows how far technology still has to go." November 2024. venturebeat.com/...-far-technology-still-has-to-go
  4. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25Glazer, E., Erdil, E., Besiroglu, T., et al. "FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI." arXiv:2411.04872, November 2024. arxiv.org/...2411.04872
  5. ^1 ^2 ^3 ^4Epoch AI. "FrontierMath Tier 4." epoch.ai/...frontiermath-tier-4
  6. ^Epoch AI. "FrontierMath: Open Problems - Hadamard Matrices." epoch.ai/...hadamard
  7. ^1 ^2OpenAI. "Announcing o3." December 20, 2024.
  8. ^1 ^2 ^3 ^4TechCrunch. "OpenAI's o3 AI model scores lower on a benchmark than the company initially implied." April 20, 2025. techcrunch.com/...an-the-company-initially-implied
  9. ^1 ^2llm-stats.com. "FrontierMath Benchmark Leaderboard." llm-stats.com/...frontiermath
  10. ^OpenAI. "Advancing science and math with GPT-5.2." openai.com/...gpt-5-2-for-science-and-math
  11. ^1 ^2Chen, E. "FrontierMath." Power Overwhelming (blog), November 10, 2024. blog.evanchen.cc/...frontiermath
  12. ^1 ^2 ^3 ^4 ^5TechCrunch. "AI benchmarking organization criticized for waiting to disclose funding from OpenAI." January 19, 2025. techcrunch.com/...-to-disclose-funding-from-openai
  13. ^1 ^2 ^3Fortune. "'Manipulative and disgraceful': OpenAI's critics seize on math benchmarking scandal." January 2025. fortune.com/...ontiermath-epoch-altman-trump-biden
  14. ^1 ^2 ^3 ^4WinBuzzer. "GPT-5.4 Pro Cracks Open Math Problem." March 24, 2026. winbuzzer.com/...blem-epoch-ai-frontiermath-xcxwbn
  15. ^1 ^2 ^3 ^4 ^5Vellum. "Everything You Need to Know About GPT-5.5." April 2026. vellum.ai/...ything-you-need-to-know-about-gpt-5-5
  16. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11Epoch AI. "FrontierMath: Tiers 1-4." May 11, 2026 update. epoch.ai/...tiers-1-4
  17. ^1 ^2 ^3 ^4 ^5Startup Fortune. "GPT-5.5 is turning AI benchmarks into an audit problem." May 2026. startupfortune.com/...hmarks-into-an-audit-problem
  18. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15Epoch AI. "FrontierMath Tiers 1-4 (v2)." June 12, 2026. epoch.ai/...about
  19. ^1 ^2 ^3 ^4 ^5 ^6 ^7DigitalApplied. "FrontierMath v2: When AI Benchmarks Get Error-Corrected." June 2026. digitalapplied.com/...rected-ai-benchmark-analysis
  20. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9OpenAI. "GPT-6 Astra: A new generation of intelligence." September 3, 2026. openai.com/...gpt-6-astra (the page was unavailable for part of the afternoon of publication; text and tables checked against contemporaneous press reports and a reader-saved copy)
  21. ^1 ^2The New Stack (Frederic Lardinois). "OpenAI launches GPT-6 Astra and says welcome to the 'AGI era'." September 3, 2026. thenewstack.io/openai-gpt6-astra-benchmarks
  22. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Epoch AI. "FrontierMath Tier 4 (v2)." AI Benchmarking Hub, data downloaded September 3, 2026. epoch.ai/...frontiermath-tier-4-v2 (data file: epoch.ai/...benchmarks.csv)
  23. ^1 ^2Epoch AI. "About FrontierMath" (conflict of interest statement and versioning). Accessed September 3, 2026. epoch.ai/...about
  24. ^1 ^2 ^3Adamczewski, T. and Burnham, G. "Announcing FrontierMath Erdős." Epoch AI, September 1, 2026. epoch.ai/...announcing-frontiermath-erdos
  25. ^VentureBeat (Carl Franzen). "'Welcome to the AGI era': OpenAI launches GPT-6 Astra." September 3, 2026. venturebeat.com/...era-openai-launches-gpt-6-astra
  26. ^1 ^2Besiroglu, T. and Sevilla, J. "Clarifying the creation and use of the FrontierMath benchmark." Epoch AI, January 23, 2025. epoch.ai/...openai-and-frontiermath
  27. ^OpenAI. "Ten advances in mathematics and theoretical computer science." August 1, 2026. openai.com/...ten-advances-in-mathematics
  28. ^OpenAI. "Improved short gaps between primes." Manuscript dated August 30, 2026. cdn.openai.com/...short_gaps.pdf
  29. ^OpenAI. "Improved long gaps between primes." Manuscript, 2026. cdn.openai.com/...long_gaps.pdf

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

8 revisions · v9 · 8,230 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independently fact-checked on September 3-4, 2026 against the cited primary sources (OpenAI, benchmark maintainers, vendor pricing pages) and press; verifier findings applied before publication.

Cite this page: AI Wiki. "FrontierMath." aiwiki.ai, updated 4 Sept 2026, fact-checked 4 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/frontiermath

Suggest edit