Home/Domains/Scientific reasoning

94.8% on GPQA Diamond.

Graduate-level questions in physics, chemistry and biology.

Clears the bar in

2027

80% interval 2026–2028

20252031

99% on GPQA Diamond, which leaves no headroom on an exam where domain experts score about 65% is the bar this date is solved against.

When GPQA Diamond reaches 99.0%

The year a frontier model first scores 99.0% on GPQA Diamond accuracy. One benchmark, not Scientific reasoning as a whole.

2025 · 0.0%
2026 · 13.0%
2027 · 48.5%
2028 · 30.9%
2029 · 6.9%
2030 · 0.7%
2031 · 0.0%
2032 · 0.0%
2033 · 0.0%
2034 · 0.0%
202520292034
Median
2027
80% interval
2026–2028
By 2030
100%

The published curve.

Treat the band as too narrow: in backtests it held 57% of later measurements, not 80%. How far to trust it

Change the assumptions

The curve redraws in your browser, labelled as yours. Nothing you set here changes the published date.

Today's reading

If you think Scientific reasoning is further along or behind than the latest measurement.

GPQA Diamond94.8%

The bar

What score on GPQA Diamond accuracy counts as done. Widen the spread if you are less sure where it sits.

Score99.0% · ±0.45

Pace of the trend

as fitted

What if progress runs faster or slower than the fitted line?

half as fasthalf again as fast

Questions

GPQA is answered by the same frontier language models Epoch's training-run series describes, so the systems measured and the systems setting the compute frontier are the same class. The reading is close to its ceiling: domain experts score about 65% and the frontier is at 94.8%, so this domain's date is as much about when the exam is exhausted as about capability, and the page says so.

No acceleration split has been justified for this series: GPQA's SOTA points sit on one line from GPT-4 onward, with no residual pattern of the kind that forced software's 2023 cut. Every SOTA measurement is fit, and this note says so.

No. Each domain is measured on its own scale - a task length in hours, an accuracy on a fixed test set - and every one of them scores a model on a bounded, pre-specified set of tasks. That is a measure of DEPTH on that set, not of breadth of competence. Depth is not breadth: a model can clear the bar on GPQA Diamond and still fail at things a child does.