Training compute is still doubling about every five months

The clearest leading indicator hasn't bent, and why it still does not move a single date on this site.

The TowardSingularity Team · Oct 2026 · 4 min read

Of the four readings this site publishes, training compute is the only one it fits a trend through. The line this site fits through Epoch AI’s figures rises 0.71 orders of magnitude a year, which has the largest training runs doubling about every five months, and the points sit close to it: the fit explains 99% of their variance.

The largest disclosed training run in the data today is Grok 4, at about 10^26.7 floating-point operations. That is roughly five hundred million billion billion operations, for one model, and the line says the frontier will be ten times past it within about eighteen months.

Why compute leads

Compute is upstream of almost everything else: more compute buys larger models, longer training, more experiments, and more attempts at the architectural lottery. It is the rare input that is measurable in advance and hard to fake. A GPU-hour cannot be quietly inflated the way a benchmark score can.

“Compute is the metronome the rest of the field marches to.”

A leading indicator is still not a guarantee. Doubling compute does not double capability, and the step from FLOP to usefulness is exactly what the capability map has to be fitted to, one domain at a time.

How the line is drawn

The source is Epoch AI’s dataset of notable models, which records an estimated training compute for every model whose developer published enough to estimate one. Three choices turn that into the line on the site.

It starts in 2010. The dataset reaches back to 1950, but one straight line across the symbolic era and the scaling era describes neither of them well.

The current year is left out of the fit. A year in progress is incomplete, and the biggest runs are often disclosed months after they finish, so a partial year reads low. 2026 shows that right now: its reading so far sits about a full order of magnitude below 2025’s. It is still plotted, marked as outside the fit, so you can watch it fill in.

And each year’s point is the average of its five largest runs, not its single largest. That one deserves a story.

The 2024 dip that wasn’t

When the site used each year’s largest run, the line had a strange kink. 2023 read 10^25.70. 2024 read 10^25.58, carried by Llama 3.1 405B. On that reading, the biggest training run of 2024 was smaller than the biggest of 2023.

Nobody believes that happened. What happened is that the largest closed models of 2024 published no specifications at all, so they have no compute estimate, and the biggest run that could be measured was an open-weights one. The maximum of a year is the maximum of whatever that year’s labs chose to disclose, and disclosure is uneven enough to swamp the signal.

Averaging the top five, after shifting each rank back onto the scale of the largest, ties a year to more than one lab’s habits. The 2023 to 2025 readings became 10^25.84, 10^26.02 and 10^26.87, the dip disappeared, and the fit tightened. The change was built to leave the overall level alone: across the sixteen fitted years the average shift is zero. It buys a steadier line, not a higher one.

Efficiency on top

Raw compute is half the picture. The same capability also gets cheaper to reach as training methods improve. Epoch estimates that separately from published algorithmic-progress evaluations, and the site fits it at about 0.44 orders of magnitude a year from nine of them. In plain terms, a given level of capability needs about 2.7 times less compute each year than the year before.

Put the two together and effective compute, the hardware and the know-how combined, grows by a little over one order of magnitude a year. That is the number to keep in mind when someone says scaling has stalled because one year’s largest run looked flat.

Three caveats before reading any of this as a forecast of capability:

  • Diminishing returns: each doubling buys less capability than the last
  • Data and energy ceilings that scaling laws ignore
  • Efficiency gains that deliver the same capability for less compute

Why it moves no date

For now the metronome has not slowed. That trend is L1 on the methodology page, and it is published as context rather than as an input to any date.

That was not always the plan. Compute used to sit at the head of a five-layer chain as if it drove each forecast. It never could. The capability layer pairs each measurement with the frontier compute level at that model’s release date, and that level is itself a straight line in the date. Once compute is a straight line in the date, a domain’s slope against compute and its slope against the calendar are the same number, and we checked that to six decimal places on all three domains. Compute cancels out of every forecast. Each date rests on its own domain’s measurements and its bar, and nothing else.

A leading indicator going dark

There is one more reason not to lean on compute, and it is getting stronger. Epoch can estimate a model’s training compute only when its developer publishes enough detail. Since 2024 that has mostly meant open-weights models. The closed frontier, the models METR actually measures, mostly publishes nothing.

When this site tried to pair each measured model with its own compute figure, exactly four of METR’s twenty-six models had one. The most watched number in AI is being read more and more from the models just behind the frontier. The trend is still clear, and it is still worth publishing. But a reader should know that each new point says a little less about the very largest systems than the points before it did.

ALL FIELD NOTES

See the data
behind
the notes.

Ten domains, each on its own measure, with the gaps published beside the measurements.