How the dates are made

Home/Methodology

Three steps, the same for every domain. Each domain’s own methodology tab carries its numbers and how far to trust them.

  1. Step 1

    Fit a line through one domain’s measurements

    Each domain with a dated series gets a straight line through its own results against the date: METR’s task lengths for software, GPQA Diamond for science, FrontierMath for maths. Task lengths go on a log scale, so a doubling is a constant step. Test scores go on a logit scale, so a score flattening near 100% is not mistaken for a slowdown.

    y(t) = ȳ + b · (t − t̄)

  2. Step 2

    Pick the bar

    A bar says what counts as done in that domain: 79 hours of unattended software work, 99% on GPQA, 90% on FrontierMath. It is a choice, not a measurement, so it is published and carries its own uncertainty. It matters less than you would think: a bar ten times harder moves software’s date by about a year.

    y* ~ N(μ, σ)

  3. Step 3

    Turn one year into a range

    Solving for when the line reaches the bar gives one year. Doing it 5,000 times, each time with the line, the bar and one model’s scatter drawn from their uncertainty, gives a range. The site publishes the median and the 80% interval, and counts years as calendar years.

    t* = t̄ + ( y* − ȳ ) / b

What could make a date wrong

  • The band is narrower than the real uncertainty.

    Refit at past cutoffs, the method’s 80% intervals have held well under 80% of the measurements that followed, because the series kept speeding up. Each domain shows its own figure.

  • A straight line may be the wrong shape.

    Every domain also shows a curved fit that lets the pace change. Where the curve clearly fits better, as it does for software, it gives an earlier date.

  • The bar can sit past what the test can measure.

    METR say task lengths above about 16 hours are unreliable on their current suite, and software’s bar is 79 hours. Confirming it will need a longer suite.

  • A score can be at its bar already.

    When a test score is within one run’s noise of its bar, the date forecasts noise rather than progress, and the domain says so. Maths is there now.

  • The live figures: Software, Science, Maths.

What is not in the date

Training compute, algorithmic efficiency, inference cost, the leaderboard rating and the prediction market are all published on the site, each in its own units with its source and date. None of them moves a date.

Compute looks as if it should. But the compute trend is itself a straight line in time, so a domain’s progress per unit of compute and its progress per year are the same number, and compute cancels out. The efficiency trend is estimated from published model evaluations, and how that is done is on its own page.

Questions

Yes, where the site reports them. Training compute comes from Epoch AI’s notable-models table and algorithmic efficiency from the evaluation set behind their paper Algorithmic Progress in Language Models, 50% task horizons on software and ML-research tasks from METR, cost from the OpenRouter price list, benchmarks from the top LMArena Elo cross-checked monthly against Stanford HELM, and the forecast from an AGI-by-2030 market. Every reading is published in the units its own source uses, with the date it was read. If a source is briefly unavailable the last good reading is held rather than a fresh one invented, and it is labelled as a held reading.

Because a single score would have to squash four readings, measured in four different units, onto one scale. That scale would be a choice and not a measurement, and the choice decides the answer. Anchor training compute between 10²³ and 10²⁸ FLOP and today reads 74. Anchor it between 10²³ and 10³⁰ and the same reading is 53. Averaging several such scales is worse still, because the width you picked for each one quietly becomes its weight. So this site reports each reading in its own units, and makes one dated claim per domain, which is a thing you can actually argue with.

One thing, per domain: the year a fitted trend reaches a bar you set, with a band around it from the uncertainty in that fit. That is a real statistical object. It has something it is estimating, a standard error, and a set of assumptions you can attack. Everything else on the site is a reading reported in its own units, with a source and a date attached, and no arithmetic done to it.

No. None of them does. Compute is fit and published, but because the effective-compute line is a straight line in the date, γ₁(β_C + β_E) equals a domain’s own slope against the calendar exactly, and compute cancels out of the solved year. The date rests on that domain’s own measurements and the bar. Cost, benchmark and market readings feed nothing either. The efficiency trend L2 is fit on published test results rather than on a price, a domain’s capability map reads that domain’s own measurement rather than a general leaderboard, and a market forecast of the answer cannot be an input to the answer without going in a circle. They are published because they are worth watching, on the home page and the domain pages, each in the units its own source uses and carrying the date it was read.

Yes, and that is a question about the domain forecasts. A domain fit on an accuracy over a fixed test set runs out of room as that test gets used up. GPQA Diamond sits near its ceiling, and the date it produces says as much about a benchmark being finished as about capability. The site flags such a domain as bar already cleared and labels the date as history rather than as a forecast, and when a score is within one test run’s noise of its bar it says that too, because the year then forecasts noise rather than capability. The fix is a harder benchmark, not a different fit, and it is why the measurement this site leads with is a task length, which has no ceiling to hit. The leaderboard rating on the home page is cross-checked monthly against Stanford HELM, and the run is flagged if the two boards drift apart, but that reading feeds no calculation either way.

Yes, and that is the design. Every domain runs the same method: a straight-line trend in its own measurements against the date, with standard errors that allow for streaky residuals, its own bar, and the same thousands of draws. The shared compute and efficiency trends are published beside it as context and as units. The equations do not change from one domain to the next, and neither does the code. A domain is a data drop. The one thing that varies is the scale a domain’s measurement is fit on. A task length with no ceiling is fit in log₁₀ hours. A percentage on a fixed test set is fit as a base-10 logit, because a percentage flattens out as it nears 100 whether or not capability flattens, and fitting the raw percentage would read that as a slowdown. Both are powers of ten, so γ₁ means the same thing everywhere and two domains can be read against each other. Because 0% and 100% sit infinitely far away on a logit scale, and real scores do hit both, every percentage metric declares its test-set size and the score is nudged off the ends by half an item. That is a real assumption, written down rather than left to be found.

Because nobody publishes a dated series for them, and the trend is the one input that cannot be invented. METR’s horizon suite covers software and ML-research tasks with no domain breakdown, and Epoch AI’s benchmarking hub supplies the science and mathematics series. For computer use, robotics, driving, video, medicine, law and finance the benchmark usually exists but the dated leaderboard does not. OSWorld renders client-side and publishes no release dates, LegalBench is a static suite, and vendor-reported scores on vendor-authored benchmarks are not a substitute. Each of those domains still says which series it is waiting on, so the gap names something specific.

By its date stamp, such as reading · 2026-08-19, taken from the page that reported it. Every pipeline run appends a row to the changelog, so a quote stays attached to the exact inputs that produced it, and nothing is back-filled or smoothed.

See it
with real
numbers.

The software forecast, with the curve, the console to move it, and every caveat beside the date.