When GPQA Diamond reaches 99.0%
The year a frontier model first scores 99.0% on GPQA Diamond accuracy. One benchmark, not Scientific reasoning as a whole.
The published curve.
Treat the band as too narrow: in backtests it held 57% of later measurements, not 80%. How far to trust it
Change the assumptions
The curve redraws in your browser, labelled as yours. Nothing you set here changes the published date.
Today's reading
If you think Scientific reasoning is further along or behind than the latest measurement.
The bar
What score on GPQA Diamond accuracy counts as done. Widen the spread if you are less sure where it sits.
Pace of the trend
as fittedWhat if progress runs faster or slower than the fitted line?
Questions
GPQA is answered by the same frontier language models Epoch's training-run series describes, so the systems measured and the systems setting the compute frontier are the same class. The reading is close to its ceiling: domain experts score about 65% and the frontier is at 94.8%, so this domain's date is as much about when the exam is exhausted as about capability, and the page says so.
No acceleration split has been justified for this series: GPQA's SOTA points sit on one line from GPT-4 onward, with no residual pattern of the kind that forced software's 2023 cut. Every SOTA measurement is fit, and this note says so.
No. Each domain is measured on its own scale - a task length in hours, an accuracy on a fixed test set - and every one of them scores a model on a bounded, pre-specified set of tasks. That is a measure of DEPTH on that set, not of breadth of competence. Depth is not breadth: a model can clear the bar on GPQA Diamond and still fail at things a child does.