The measurement
Each dot is one model's 50% task horizon, plotted on its release date. Nothing on this chart is a forecast.
- Doubling
- ~4.1 mo
- 125 days · 2023 onward
- Fit
- r² 0.94
- 15 of 18 measurements fitted
- Measured span
- 3 sec 17 h
- Feb 2019 → Apr 2026
How to read this chart
Our fit over 2023 onward doubles every 125 days. METR publish 129 days over the same window, and 188 all-time against our 183. We reproduce their number rather than assert our own.
Measurements
18 modelsEvery point on the chart above, named. The date is the model's release, which is what the trend is regressed on.
| Model | Released | Position on the measured range | Task horizon |
|---|---|---|---|
| claude-mythos-preview-early | Apr 2026 | 17 h | |
| claude-opus-4-6 | Feb 2026 | 12 h | |
| gpt-5-2 | Dec 2025 | 5.9 h | |
| claude-opus-4-5 | Nov 2025 | 4.9 h | |
| gemini-3-pro | Nov 2025 | 3.7 h | |
| gpt-5-2025-08-07 | Aug 2025 | 3.4 h | |
| o3 | Apr 2025 | 2.0 h | |
| claude-3-7-sonnet | Feb 2025 | 1.0 h |
Questions
METR times the same frontier language models that Epoch's largest-training-run series describes, so extrapolating along that line is a claim about the same systems. It is still not evidence that compute causes capability: the frontier level is a function of the date, so this fit is a horizon-vs-time trend expressed in compute units.
METR publish a doubling time of 187.8 days all-time against 128.7 days from 2023 on, and 2023 is where they split it. Fit on everything and the last six points sit above the line by a growing margin, ending +0.51 dex out.
No. Each domain is measured on its own scale - a task length in hours, an accuracy on a fixed test set - and every one of them scores a model on a bounded, pre-specified set of tasks. That is a measure of DEPTH on that set, not of breadth of competence. Depth is not breadth: a model can hold a 79-hour Software & ML research job unattended and still fail at things a child does.