Home/Domains/Advanced mathematics

87.8% on FrontierMath T4.

Research-level problems set by working mathematicians, and contest problems below them.

Clears the bar in

2026

80% interval 2026–2026

20252028

90% of FrontierMath Tier 4: nine in ten research-level problems, each a specialist's day of work is the bar this date is solved against.

⚠ FrontierMath Tier 4 accuracy needs 37 of 41 items and the latest reading has 36: 1 more. One run of the same model moves by about ±2 items on a test this size, so this domain is at its bar within measurement noise, and the year below is a forecast of that noise rather than of new capability.

The measurement

Each dot is one model's FrontierMath Tier 4 accuracy, plotted on its release date. Nothing on this chart is a forecast.

Accuracy on FrontierMath Tier 4Epoch AI · Benchmarking Hub ↗
-0.6%1.5%10.5%37.9%75.3%94.5%202520262027o3-mini · -0.0% · Jan 2025o4-mini · 4.9% · Apr 2025GPT-5 · 22.0% · Aug 2025GPT-5.2 Pro · 46.0% · Dec 2025GPT-5.4 · 49.0% · Mar 2026GPT-5.4 Pro · 58.5% · Mar 2026GPT-5.5 Pro · 78.0% · Apr 2026Claude Fable 5 · 87.8% · Jun 2026Latest87.8%Claude Fable 5 · Jun 2026
Measured, before the fit windowMeasured (Epoch AI · Benchmarking Hub)Fitted trend · the rows in the fit window
Doubling
~2.0 mo
62 days · all measurements
Fit
r² 0.96
8 of 8 measurements fitted
Measured span
-0.0% 87.8%
Jan 2025 → Jun 2026
How to read this chart

The axis is a log-odds scale, the same one the trend is fit on. Equal distance means an equal cut in the remaining error, so 50% to 90% is about the same step as 90% to 99%, which is why the labels crowd near the top. A percentage axis would flatten near 100% whether or not capability flattened. Very close to the ceiling the steps shorten, because a 41-question test cannot resolve past its own last item.

Our fit over all measurements doubles every 62 days. Epoch AI · Benchmarking Hub publish the scores and the release dates. The trend line through them is ours, and it is the only thing on this chart that is.

Measurements

8 models

Every point on the chart above, named. The date is the model's release, which is what the trend is regressed on.

Advanced mathematics measurements, most recent first
Model Released Position on the measured rangeFrontierMath T4
Claude Fable 5Jun 202687.8%
GPT-5.5 ProApr 202678.0%
GPT-5.4Mar 202649.0%
GPT-5.4 ProMar 202658.5%
GPT-5.2 ProDec 202546.0%
GPT-5Aug 202522.0%
o4-miniApr 20254.9%
o3-miniJan 20250.0%

Questions

FrontierMath is answered by the same frontier language models Epoch's largest-training-run series describes, so the systems measured and the systems setting the compute frontier are the same class. Contest sets such as Mock AIME are solved outright by frontier models, so they have no signal left to fit, which is why this domain is measured on the research-level tier.

The whole Tier 4 series is post-2025, because it did not exist earlier, so there is no earlier era to split off. Every SOTA measurement is fit.

No. Each domain is measured on its own scale - a task length in hours, an accuracy on a fixed test set - and every one of them scores a model on a bounded, pre-specified set of tasks. That is a measure of DEPTH on that set, not of breadth of competence. Depth is not breadth: a model can clear the bar on FrontierMath T4 and still fail at things a child does.