87.8% on FrontierMath T4.
Research-level problems set by working mathematicians, and contest problems below them.
Clears the bar in
2026
80% interval 2026–2026
20252028
90% of FrontierMath Tier 4: nine in ten research-level problems, each a specialist's day of work is the bar this date is solved against.
⚠ FrontierMath Tier 4 accuracy needs 37 of 41 items and the latest reading has 36: 1 more. One run of the same model moves by about ±2 items on a test this size, so this domain is at its bar within measurement noise, and the year below is a forecast of that noise rather than of new capability.
How this date is made
- A straight line through 8 measurements of FrontierMath Tier 4 accuracy, on a logit scale. It rises 1.79 on that scale each year, and fits the points with r² 0.96.
- The bar is 90.0%, 90% of FrontierMath Tier 4: nine in ten research-level problems, each a specialist's day of work. It is a choice, not a measurement: a bar that cuts the remaining error about tenfold costs about 7 months.
- Solving for when the line reaches the bar, 5,000 times with the line, the bar and a single model's scatter drawn from their uncertainty, gives 2026, with an 80% interval of 2026–2026.
Latest measurement 87.8% (Claude Fable 5). Source: Epoch AI · Benchmarking Hub.
How far to trust it
- FrontierMath Tier 4 accuracy needs 37 of 41 items and the latest reading has 36: 1 more. One run of the same model moves by about ±2 items on a test this size, so this domain is at its bar within measurement noise, and the year below is a forecast of that noise rather than of new capability.
- The band has held up so far. Refit at 3 past cutoffs, its 80% intervals held 100% of the measurements that followed, though 3 is too few to say much.
- No clear change of pace. A curved fit does not beat the straight line here.
- It is an extrapolation. The bar is 0.09 steps on the fitted scale past the latest reading, about 1 month at the fitted pace.
- 41% of the bar's uncertainty range sits at or below today's reading, so it is already reached; the date covers the rest.
The numbers and equations
- y(t) = ȳ + b · (t − t̄)
- y* ~ N(μH*, σH*), above today's reading
- t* = t̄ + ( y* − ȳ ) / b
A draw that would have crossed before the latest measurement, which did not, is dropped (8.0% here). Years are calendar years. The whole Tier 4 series is post-2025, because it did not exist earlier, so there is no earlier era to split off. Every SOTA measurement is fit.
Training compute and efficiency (β_C 0.715, β_E 0.438 a year) are fit and published, but γ₁(β_C + β_E) equals b exactly, so they cancel out of the date.
| b | 1.788 | slope: change in y per calendar year |
| se(b) | 0.148 | its standard error, widened 1.00× for residuals that run in streaks |
| σ | 0.195 | how far a single model sits off the line |
| μH* | 0.91 | the bar on the fitted scale (90.0%) |
| σH* | 0.40 | uncertainty on the bar |
| γ1 | 1.550 | the same slope in compute units, for context |
Questions
FrontierMath is answered by the same frontier language models Epoch's largest-training-run series describes, so the systems measured and the systems setting the compute frontier are the same class. Contest sets such as Mock AIME are solved outright by frontier models, so they have no signal left to fit, which is why this domain is measured on the research-level tier.
The whole Tier 4 series is post-2025, because it did not exist earlier, so there is no earlier era to split off. Every SOTA measurement is fit.
No. Each domain is measured on its own scale - a task length in hours, an accuracy on a fixed test set - and every one of them scores a model on a bounded, pre-specified set of tasks. That is a measure of DEPTH on that set, not of breadth of competence. Depth is not breadth: a model can clear the bar on FrontierMath T4 and still fail at things a child does.