Field Notes/Methodology

What these forecasts do not measure

The dated forecasts are not about AGI. Each one is about a single measurement in a single domain. What those numbers cover, and why raising the bar does not widen them.

The TowardSingularity Team · Oct 2026 · 5 min read

This site publishes a dated number only for a domain whose data supports one, and the most common way to misread it is to treat it as an AGI date. It is not one. This post is about what those numbers actually cover, which is narrower than the label people reach for, and more useful.

What each number actually is

Each dated forecast rests on one measurement in one domain. For software and ML research it is METR’s 50% time horizon: the length of task, measured by how long a human professional takes, that a model finishes on its own about half the time. For scientific reasoning it is the score on GPQA Diamond, a fixed graduate-level exam, and for mathematics it is the score on FrontierMath Tier 4, a set of research-level problems.

So each model answers a real and reasonably precise question, such as: when will a frontier model finish a two-working-week software task unattended, right about half the time? None of them answers, or has any way to answer, whether that same model could assemble flat-pack furniture, make coffee in an unfamiliar kitchen, or carry what it knows into a domain it has never seen. Those are the things most definitions of AGI actually require.

“Depth is not breadth. A model can hold an 80-hour coding job, or ace a research exam, and still fail at things a child does.”
The scope note carried on every domain page

Ten domains, three dates

The site tracks ten domains. Three of them publish a date: software, scientific reasoning and mathematics. The other seven do not, and the reason is never that we decided they did not matter.

Six of the seven are unmeasured. A benchmark usually exists, but nobody publishes it as a dated series of frontier results, and a trend cannot be fitted through points that were never recorded. Each of those domain pages names the series it is waiting on, so the gap points at something specific:

  • Agentic computer use needs a dated OSWorld or WebArena series
  • Robotic manipulation needs RLBench horizons with human baselines
  • Self-driving has only a community telemetry tracker, which is not a controlled benchmark
  • Medicine, law and finance each lack a published human baseline for how long the work takes

The seventh, video, is a different case. The obvious proxy is clip length, and METR flag that it correlates weakly with difficulty. A two-hour film is not a two-hour task. There the instrument is wrong, not merely missing, so the domain is marked as one where this method does not apply.

The gaps are not spread evenly, either. The domains with dates are the ones where models work on text in a sandbox. The domains without them are the ones that touch the physical world, or institutions, or both. Any reading of three dates as a general timeline quietly assumes the other seven move at the same pace, and the scraps of evidence we do have say they do not.

The domain most likely to set the floor

Computer use is the clearest example. It is driven by the same frontier language models as software, scaffolded differently, so it is not a different class of system. The horizons reported for it are 40 to 100 times shorter than software’s, improving at a similar rate.

If that holds, a model that can carry a two-week coding project still could not reliably book a trip through a handful of ordinary websites. For most practical definitions of general capability, that is the binding constraint, and it is the one the site cannot date yet. That is why the domain page calls it the gap most likely to set an AGI floor.

Self-driving is the starker case. The figure reported for it is about 0.6 doublings a year, roughly five times slower than software. Robotics is described as improving far slower than the software cluster. Neither comes from a series clean enough to fit, so neither gets a date, but both are a warning against reading the fastest curve as the general one.

Why raising the bar does not fix it

The obvious repair is to demand more: keep the AGI label and set the bar higher. It does not work, and the model says so itself. Because the fitted trends are steep, a bar ten times harder costs somewhere between half a year and two years of arrival date, depending on the domain. On software, going from two working weeks to about four person-years, a hundred times harder, buys a little over two.

You can work that through on each domain’s forecast page and move the bar yourself. The insensitivity is the finding. Depth on a fixed set of tasks and breadth are different axes, and only the first is in the model, so no bar on it can stand in for the second.

How to read a date that is here

Take software. The published median is 2027, with an 80% band from 2026 to 2028. Spelled out in full, that says: if METR’s 50% horizon keeps rising along the line fitted to its measurements since 2023, a frontier model finishes a task that takes a professional about 79 hours, unattended, half the time, most likely in 2027.

Every qualifier in that sentence is doing work. The 79 hours is a choice, roughly two working weeks, and the bar is 4.6 times beyond the latest reading of 17 hours, so the date is an extrapolation and not a reading of anything already measured. METR itself cautions that horizons above about 16 hours are unreliable on its current task suite, which means the bar could not be confirmed with today’s tasks even if a model cleared it.

And the band is narrower than the real uncertainty. Refitting the same method at past cutoffs and checking it against what came next, the 80% band caught the later measurements 29% of the time on software and 57% on science. The series have been speeding up, and a straight line under-calls a series that is speeding up. The forecast page says this beside the chart rather than in a footnote.

What the site does about it

Three rules follow from all of that, and all three are visible on the site:

  • Each forecast is labelled by what it measures: a target task length or a target score, not an AGI bar
  • Every mention of METR says 50%, and says software
  • The coverage board publishes the domains nobody has a measurement for, rather than quietly averaging around them

None of this changes the numbers. Each one was always an extrapolation of one measurement in one domain. Saying so plainly is the whole difference between a forecast you can check and a headline you can only believe.

ALL FIELD NOTES

See the data
behind
the notes.

Ten domains, each on its own measure, with the gaps published beside the measurements.