Field Notes/Methodology

Why the forecast gives no single date

A point estimate hides everything that matters. The case for publishing a distribution, and how this one is built.

The TowardSingularity Team · Oct 2026 · 4 min read

A single number with a date attached feels authoritative. It is also, almost always, wrong. Worse, it quietly discards the one thing a careful reader needs most: how uncertain the estimate actually is.

A point estimate is a compression artefact

When someone says “AGI by 2030,” they have taken a wide, lumpy probability distribution over many possible years and crushed it into a single token. The confidence interval, the part that separates planning from panicking, evaporates in the compression.

“A date is a marketing object. A distribution is an argument that can be checked.”

So the site does not lead with a date. It publishes each reading in the units its own source uses, and separately, for each domain whose data supports one, a probability curve over the year that domain’s fitted trend reaches its bar. The median is one point on that curve. The spread around it carries the rest of the information.

What the distribution preserves that a bare date throws away:

  • The 80% interval: how wide the plausible window really is
  • How sensitive the answer is to the assumptions behind it
  • Where reasonable, informed people are allowed to disagree
  • Which new evidence would actually move the curve

Where the spread comes from

Each curve is 5,000 simulated futures. Every one of them draws three things, and each draw stands for a different kind of doubt.

The first is the trend itself. A line fitted through fifteen or eighteen points is pinned down only so well, so each draw takes a slightly different level and slope, in proportion to how uncertain the fit is. Frontier results also arrive in streaks: a model above the line tends to be followed by another one above it. Ordinary regression treats every point as independent and so claims more certainty than it has. The fit corrects for that, and the correction is not small. It widens the slope’s error 1.54 times on software and 1.40 times on science.

The second is scatter. Even if the trend were known exactly, a single system lands above or below it. Each draw adds that scatter at the moment of crossing.

The third is the bar, which is a choice. Two working weeks of software, or 99% on a hard exam, is a judgement about what counts, and the draw treats it as one, with a spread around the stated value. One detail matters here. The bar is drawn only above today’s reading, because a draw that puts it below where models already are is not a forecast of anything. Those draws used to land as early arrivals and quietly pulled the low end of the band forward.

Draws that the latest measurement contradicts are also dropped. If a simulated trend reaches the bar before the most recent result, and that result did not reach it, the draw is ruled out by the evidence.

What one number hides, in practice

Two small episodes from this site make the case better than any argument.

The first is rounding. The pipeline used to turn each simulated crossing into a year by rounding the decimal, so a crossing at 2027.6, which is August 2027, was counted towards 2028. Fixing that moved science’s published median from 2028 to 2027 and maths’ from 2027 to 2026. The curves barely changed. Anyone quoting the single number saw a whole year vanish; anyone reading the curve saw almost nothing happen, which was the truth.

The second is mathematics. Its bar is 90% on FrontierMath Tier 4, which on a 41-problem test means 37 correct. The latest reading is 36. One run of the same model on a test that size moves by about two problems either way, so the bar is inside the noise. The point estimate says 2026 with complete assurance. The distribution shows that 41% of draws put the bar at or below where models already are, and the page flags the domain as at its bar. A single date cannot carry that sentence. A curve can.

What the band does not carry

The 80% band covers the three sources above. It does not cover the possibility that a straight line is the wrong shape, and that turns out to be the largest term of all.

We check this by refitting the method at every past cutoff and asking whether the band would have caught what came next. On software it did 29% of the time, against the 80% it claims. On science, 57%. Each refit has also pulled the crossing earlier: software’s moved from 2030.5 with data up to late 2024 to 2027.3 with data up to early 2026. The series have been speeding up, and a straight line under-calls a series that is speeding up every time.

The site also fits a curved trend alongside the straight one. On software the curvature is clearly positive, and the curved fit puts the median a year earlier, in 2026. On science and maths the curvature is within its noise, and the page says so. None of this is hidden. It is printed beside the chart, because a band presented without its own track record is just a confident date with extra decoration.

Disagreeing with it

The inputs are there to be disagreed with. That is the point. Move the trend’s pace or the bar on a domain’s forecast page and the arrival curve re-runs on the spot, using the same fitted model that produced the published one. If you think two working weeks is too low a bar for software, set it to two months and watch what happens to the date. If you think the recent speed-up will not last, slow the pace down.

A date can only be accepted or rejected. A distribution can be interrogated.

ALL FIELD NOTES

See the data
behind
the notes.

Ten domains, each on its own measure, with the gaps published beside the measurements.