Every domain, and where it stands.

Ten areas of cognitive work. Three have a dated series to fit, and get a forecast. The other seven do not, and each one says why.

Domain coverage

3 of 10 domains carry a date

Writing, debugging and shipping code; running machine-learning experiments.

Today
17 h
Clears
2027
Task horizonBar 79 h

Graduate-level questions in physics, chemistry and biology.

Today
94.8%
Clears
2027
GPQA DiamondBar 99.0%

Research-level problems set by working mathematicians, and contest problems below them.

Today
87.8%
Clears
2026
FrontierMath T4Bar 90.0%

Within one test run's noise of the bar.

The other 7

A benchmark exists, but no dated series

Someone measures these, but not in a form a trend can be fitted to yet.

  • Agentic computer use

    Driving a real desktop or browser to finish an errand end to end.

    Would be measured on OSWorld success rate, from OSWorld leaderboard. Needs a dated OSWorld / WebArena series. This is the gap most likely to set an AGI floor.

    Reported elsewhere: Horizons 40–100× shorter than software, improving at a similar rate.

  • Robotic manipulation

    Physical tasks in the world: the coffee test, the flat-pack test.

    Would be measured on RLBench success rate, from RLBench. Needs a dated RLBench horizon series with human baselines.

    Reported elsewhere: Improving far slower than the software cluster.

  • Self-driving

    Kilometres of unsupervised driving between human interventions.

    Would be measured on miles between disengagements, from California DMV disengagement reports. Needs a defensible series; the public tracker is community telemetry, not a controlled benchmark.

    Reported elsewhere: About 0.6 doublings per year, roughly five times slower than software.

Nothing anyone maintains

No independent, dated measurement exists, and this site will not invent one.

  • Medicine

    Diagnosis, treatment planning, and carrying a case over time.

    Clinical work carries over weeks and the outcome is often unobservable at the horizon of the task, so "a doctor takes N hours" is not a well-posed baseline for most of it. Existing medical benchmarks score answer accuracy on vignettes, which is a different quantity from how long a case a system can carry unattended.

  • Law

    Research, drafting, and running a matter to a conclusion.

    Billable-hour records exist in enormous volume but are confidential and are not task-decomposed, so the one dataset that would give human baselines cheaply is the one nobody can publish. Public legal benchmarks test retrieval and classification, not how long a matter a system can run.

  • Finance

    Analysis, modelling, and decisions carried over a horizon.

    The work whose value is easiest to measure is also the work whose results are least likely to be shared, and success is confounded with market luck over exactly the horizons that would need measuring. Timed professional baselines exist inside firms and are not published.

The wrong instrument

Data exists, but it measures the wrong thing for this question.

  • Video understanding

    Following and reasoning about long-form footage.

    Video length is used as a proxy for task length, and METR flag that the correlation with difficulty is weak. A two-hour film is not a two-hour task. The instrument is wrong here, not merely missing.

How to read the board

  • The gate

    A domain gets a date when it has a dated series and a declared bar, and not otherwise.

    The gate is the data, not a judgement about the domain.

  • The track

    Distance to the bar is drawn on the scale the fit is run in: log₁₀ hours, or a base-10 logit for an accuracy.

    On that scale a doubling is a constant step, which is why the last stretch of a track is the longest part of the climb.

  • Scope

    This board does not forecast AGI.

    Every domain is scored on a bounded, pre-specified set of tasks: depth on that set, not breadth of competence. A date says when one measurement reaches one bar, nothing more.