How this date is made
- A straight line through 15 measurements of 50% task horizon since 2023, on a log scale. It doubles about every 4.1 months, and fits the points with r² 0.94.
- The bar is 79 h, roughly two working weeks, unattended. It is a choice, not a measurement: a task ten times longer costs about 1.1 years.
- Solving for when the line reaches the bar, 5,000 times with the line, the bar and a single model's scatter drawn from their uncertainty, gives 2027, with an 80% interval of 2026–2028.
Latest measurement 17 h (claude_mythos_preview_early_inspect). Source: METR · METR-Horizon-v1.1.
How far to trust it
- The band is too narrow. Refit at 10 past cutoffs, its 80% intervals held 29% of the measurements that followed.
- The pace has been speeding up. A curved fit that allows for it gives 2026 (80%: 2026–2027).
- It is an extrapolation. The bar is 4.6× past the latest reading, about 9 months at the fitted pace. METR caution that horizons above about 16 hours are unreliable on their current task suite, and the latest reading is already 17 h. A 79 h bar is about five times past what that suite can measure, so confirming it will need a longer suite, and METR's revision from version 1.0 to 1.1 moved recent readings by up to 20%.
- 7% of the bar's uncertainty range sits at or below today's reading, so it is already reached; the date covers the rest.
The numbers and equations
- y(t) = ȳ + b · (t − t̄)
- y* ~ N(μH*, σH*), above today's reading
- t* = t̄ + ( y* − ȳ ) / b
A draw that would have crossed before the latest measurement, which did not, is dropped (0.9% here). Years are calendar years. METR publish a doubling time of 187.8 days all-time against 128.7 days from 2023 on, and 2023 is where they split it. Fit on everything and the last six points sit above the line by a growing margin, ending +0.51 dex out.
Training compute and efficiency (β_C 0.715, β_E 0.438 a year) are fit and published, but γ₁(β_C + β_E) equals b exactly, so they cancel out of the date.
| b | 0.882 | slope: change in y per calendar year |
| se(b) | 0.095 | its standard error, widened 1.54× for residuals that run in streaks |
| σ | 0.205 | how far a single model sits off the line |
| μH* | 1.90 | the bar on the fitted scale (79 h) |
| σH* | 0.45 | uncertainty on the bar |
| γ1 | 0.765 | the same slope in compute units, for context |
Questions
METR times the same frontier language models that Epoch's largest-training-run series describes, so extrapolating along that line is a claim about the same systems. It is still not evidence that compute causes capability: the frontier level is a function of the date, so this fit is a horizon-vs-time trend expressed in compute units.
METR publish a doubling time of 187.8 days all-time against 128.7 days from 2023 on, and 2023 is where they split it. Fit on everything and the last six points sit above the line by a growing margin, ending +0.51 dex out.
No. Each domain is measured on its own scale - a task length in hours, an accuracy on a fixed test set - and every one of them scores a model on a bounded, pre-specified set of tasks. That is a measure of DEPTH on that set, not of breadth of competence. Depth is not breadth: a model can hold a 79-hour Software & ML research job unattended and still fail at things a child does.