Getting the same result, for less.

Home/Methodology/Efficiency

AI gets better partly because companies spend more on training, and partly because the methods themselves improve. This page separates the two, using 208 published test results, and shows every step of the arithmetic. The number is published as context: no date on this site depends on it, as the methodology explains.

βE
0.438
powers of ten a year
in plain terms
2.74×
cheaper every year
cost halves every
8.2
months
worked out from
208
real results, 2015 to 2023

Two things make AI better. We want to measure one of them.

Suppose a model this year beats one from three years ago. Some of that is simply money: it was trained on far more computing power. The rest is skill, meaning people found better ways to build and train it. Only the second one is efficiency.

To measure it, you have to hold the money still and see what is left. That is the whole idea, and the picture beside this is what it looks like.

Each dot is one model, tested once. Left to right is how much compute trained it. Up and down is its score, and lower is better.

  • Dots fall as you go right, because more compute buys a better score.
  • The two lines are the same relationship in 2018 and in 2023. The later one sits lower everywhere.
  • That drop, marked by the upright bar between the two lines, happened without extra compute. It is what we are trying to put a number on.

Every result on the WT103 test

score down, compute across
1018102010221024710152030QRNN · 2018.1 · trained on ten to the 17.6 · scored 334 layer QRNN (h=2500) · 2018.2 · trained on ten to the 17.4 · scored 33LSTM (Hebbian, Cache, MbPA) · 2018.2 · trained on ten to the 19.4 · scored 29.2Transformer (Adaptive Input Embeddings) · 2018.7 · trained on ten to the 18.9 · scored 18.7TrellisNet · 2018.8 · trained on ten to the 18.4 · scored 29.19TrellisNet-MoS (1.4x larger) · 2018.8 · trained on ten to the 18.4 · scored 29.19Transformer-XL Large · 2019.0 · trained on ten to the 19.0 · scored 18.3GPT-2 (1542M) · 2019.1 · trained on ten to the 21.2 · scored 17.48GPT-2 (762M) · 2019.1 · trained on ten to the 20.9 · scored 22.05GPT-2 (345M) · 2019.1 · trained on ten to the 20.5 · scored 26.37Transformer-XL Large + Phrase Induction · 2019.4 · trained on ten to the 17.2 · scored 17.4AdvSoft + 4 layer QRNN + dynamic evaluation · 2019.4 · trained on ten to the 17.6 · scored 284 layer QRNN + dynamic evaluation · 2019.4 · trained on ten to the 17.6 · scored 31.6Tensorized Transformer (257M) · 2019.5 · trained on ten to the 18.7 · scored 21.2Tensorized Transformer (core-2) · 2019.5 · trained on ten to the 18.2 · scored 18.9Tensorized Transformer (151M) · 2019.5 · trained on ten to the 18.4 · scored 18.8All-attention network + adaptive span · 2019.5 · trained on ten to the 19.7 · scored 20.6DEQ-Transformer (Medium, Adaptive Embedding) · 2019.7 · trained on ten to the 17.9 · scored 23.2Megatron-LM (8.3B) · 2019.7 · trained on ten to the 22.0 · scored 10.81Megatron-LM (355M) · 2019.7 · trained on ten to the 20.6 · scored 19.31Sandwich Transformer · 2019.9 · trained on ten to the 20.2 · scored 17.84Compressive Transformers for Long-Range Sequence Modelling · 2019.9 · trained on ten to the 20.2 · scored 17.1Transformer-XL DeFINE (107M) · 2019.9 · trained on ten to the 18.7 · scored 25.72Adaptive LSTM + DeFINE · 2019.9 · trained on ten to the 18.8 · scored 35.94Transformer-XL DeFINE (141M) · 2019.9 · trained on ten to the 18.8 · scored 24.17TaLK Convolution · 2020.1 · trained on ten to the 19.4 · scored 23.3Turing-NLG · 2020.1 · trained on ten to the 22.2 · scored 10.21Feedback Transformer · 2020.1 · trained on ten to the 19.3 · scored 18.3TransformerXL + spectrum control · 2020.2 · trained on ten to the 17.7 · scored 23.2Tensor-Transformer(1core)+PN (WT103) · 2020.2 · trained on ten to the 18.2 · scored 17.9Segatron XL base, M=384 · 2020.3 · trained on ten to the 18.2 · scored 22.5Segatron XL large, M=384 · 2020.3 · trained on ten to the 19.4 · scored 17.1GPT3-6.7B (rerun of original) · 2020.4 · trained on ten to the 22.1 · scored 9.136-Layer-Tensor-Transformer+AdaHessian · 2020.4 · trained on ten to the 18.2 · scored 19.9DeLight · 2020.6 · trained on ten to the 19.4 · scored 24.14Transformer+Recurrent Windows of Context · 2020.6 · trained on ten to the 20.1 · scored 26.73Shortformer · 2021.0 · trained on ten to the 18.5 · scored 18.15ERNIE-Doc (151M) · 2021.0 · trained on ten to the 19.3 · scored 21ERNIE-Doc (247M) · 2021.0 · trained on ten to the 19.5 · scored 16.8Subformer (83M) · 2021.0 · trained on ten to the 18.6 · scored 20.88Subformer (122M) · 2021.0 · trained on ten to the 18.7 · scored 19.9Subformer (96M) · 2021.0 · trained on ten to the 18.6 · scored 20.39Linear Transformer (large) · 2021.1 · trained on ten to the 18.6 · scored 31.5Linear Transformer (small) · 2021.1 · trained on ten to the 18.5 · scored 35.5SRU++ Large · 2021.2 · trained on ten to the 19.0 · scored 17.1SRU++ Large only 2 attention layers (k=5) · 2021.2 · trained on ten to the 18.9 · scored 17.3SRU++ Base · 2021.2 · trained on ten to the 18.8 · scored 18.3RFA-GATE-Gaussian-Stateful Big · 2021.2 · trained on ten to the 18.9 · scored 23.5GLM-10B-bidirectional · 2021.2 · trained on ten to the 22.6 · scored 11.33GLM-10B-unidirectional · 2021.2 · trained on ten to the 22.6 · scored 12.22Transformer-C · 2021.3 · trained on ten to the 18.3 · scored 25.1Delta RNN (+ full context) · 2021.4 · trained on ten to the 18.0 · scored 32.8DEQ-Transformer (Post-LN) + Jacobian Regularisation · 2021.5 · trained on ten to the 19.5 · scored 24.9Adaptive Input Transformer + RD · 2021.5 · trained on ten to the 19.9 · scored 18.07GPT-2 (1.5B, Curriculum Learning 45K) · 2021.6 · trained on ten to the 20.8 · scored 13.72ALiBi (L=3072, Lvalid = 3072) · 2021.7 · trained on ten to the 20.3 · scored 18.3$\infty$-former (SM) · 2021.7 · trained on ten to the 20.1 · scored 16.61PermuteFormer · 2021.7 · trained on ten to the 18.5 · scored 32.49base LM+GNN+kNN · 2021.8 · trained on ten to the 18.9 · scored 16.8S4 · 2021.8 · trained on ten to the 19.9 · scored 20.95GPT2+CoreLM+Fine-Tuning · 2021.8 · trained on ten to the 16.5 · scored 29.51GPT-2-Medium+Pixelfly · 2021.9 · trained on ten to the 19.1 · scored 21GPT-2-Small+Pixelfly · 2021.9 · trained on ten to the 18.6 · scored 22.5Gopher (7.1B) · 2021.9 · trained on ten to the 23.8 · scored 10.81Gopher (280B) · 2021.9 · trained on ten to the 22.1 · scored 8.12HSO · 2022.0 · trained on ten to the 20.5 · scored 20.3GPT3-6.7B + muP · 2022.2 · trained on ten to the 22.1 · scored 8.56Segatron-XL large, M=384 + HCP · 2022.2 · trained on ten to the 19.4 · scored 17Transformer Large + HCP · 2022.2 · trained on ten to the 18.8 · scored 25.3Segatron -XL base, M=150 + HCP · 2022.2 · trained on ten to the 18.2 · scored 22.1MemSizer · 2022.2 · trained on ten to the 18.9 · scored 20.8Chinchilla · 2022.2 · trained on ten to the 23.8 · scored 7.16NoPos · 2022.2 · trained on ten to the 20.2 · scored 20.97Monarch-GPT-2-Medium · 2022.3 · trained on ten to the 20.6 · scored 20.3Monarch-GPT-2-Small · 2022.3 · trained on ten to the 20.3 · scored 20.7LaMemo · 2022.3 · trained on ten to the 18.9 · scored 23.77B2T connection (16L) · 2022.4 · trained on ten to the 19.4 · scored 19.2DITTO · 2022.4 · trained on ten to the 19.0 · scored 24.33NMST+GPT-2 · 2022.8 · trained on ten to the 20.1 · scored 20.69Decaying Fast Weights Transformer · 2022.8 · trained on ten to the 19.1 · scored 20.5Transformer + GFM · 2022.9 · trained on ten to the 18.9 · scored 20.05Hybrid H3-355M · 2023.0 · trained on ten to the 20.1 · scored 16.9Hybrid H3-125M · 2023.0 · trained on ten to the 19.6 · scored 23.7Hybrid H3-2.7B · 2023.0 · trained on ten to the 20.9 · scored 10.6Hybrid H3-1.3B · 2023.0 · trained on ten to the 20.6 · scored 12.5Sparse Wide GPT-3 Small · 2023.2 · trained on ten to the 19.9 · scored 20.4
20182023 the drop compute did not buy

Wait a year and the same result costs 2.74 times less

0.438 is a power of ten, not a multiplier, so turn it back into one: 100.438 = 2.74. A job that needs a million pounds of computing today needs about 364 thousand next year, and about 133 thousand the year after, for the same result.

Put another way, the price of any fixed level of ability halves roughly every 8.2 months. Halving means dividing by 2, and 2 is 100.3010, so the question is how long 0.438 powers of ten a year takes to add up to 0.3010 of them:

0.3010 ÷ 0.438 = 0.6867 years

0.6867 × 12 = 8.2 months

That is faster than computer chips have ever got cheaper on their own, and it is why capability spreads outward so quickly: what only the largest labs could afford three years ago is ordinary now.

Two real models, same test

2015.2genCNN + dyn eval
score 106.3 · trained on 1016.86 of compute
2019.9bRSM + cache
score 103.5 · trained on 1014.33 of compute

16.86 − 14.33 = 2.53 powers of ten

102.53 = 339× less compute

The later one scored a shade better on that much less compute, 4.7 years on. It was picked by rule, not by eye: same test, at least three years apart, the later score no worse, and of every pair meeting that, the widest compute gap. One pair proves nothing by itself. The number above is the same thing averaged over all 208.

Where this goes next: the forecast adds this rate to the rate at which training runs are getting bigger, and calls the total "effective compute". A domain's own measurements are then fitted against that, and the year the fitted line reaches the bar you set is the date the site publishes.

Show the full working

The equation, all 208 results, how they are prepared, the arithmetic step by step, and how sure the number is.

One line that says what a score is made of

A model's score comes from three things: how hard the test is, how much compute it was trained on, and what the field had learned by then. Written down, that is:

log10 P = atest + b * log10 C + gyear

Scores and compute are both written as powers of ten, because the numbers involved run from thousands to trillions and nothing else fits on one page. A compute of 20 means 1020, a one followed by twenty zeros.

P
the score a model got on a set test. Lower is better here, like a golf score. Its technical name is perplexity.
C
how much calculation went into training that model.
b
what ten times the compute is worth. One number, shared by every row.
atest
how hard each test is. Three tests are used and they are not equally hard.
gyear
whatever that year did better that the compute does not explain. This is the part we are after.

We know the scores and we know the compute, because both are published. The two unknowns are b and each year's own gain gyear, and the computer finds the values for them that come closest to matching all 208 results at once. Everything below is what came back.

Every result the answer was built from

208 results covering 202 models, on 3 standard tests, from 2015 to 2023. Nothing is weighted or held back. The table beside this is the entire input.

These are the rows that survived the filtering, which the next section walks through row by row.

What a row has to contain

system
which model was tested
year
when it came out
flop
the compute its training used
benchmark
which test the score is from
perplexity
the score itself. The column keeps its technical name, which is what this kind of score is called.

Nothing here is stored or typed in by hand. The whole page is worked out again on every update, so new results move the answer on their own.

208 rows, newest first
Every evaluation the fit was built from, newest first
modelyearcomputetestscore
LLaMA-33B (LoRA finetuned)2023.3923.48PTB7.68
LLaMA-13B (LoRA finetuned)2023.3922.93PTB8.64
LLaMA-7B (LoRA finetuned)2023.3922.66PTB9.69
LLaMA-13B (LoRA finetuned)2023.3922.93WT25.54
LLaMA-65B (LoRA finetuned)2023.3923.78WT24.27
LLaMA-7B (LoRA finetuned)2023.3922.66WT26.19
MPT-7B2023.3422.62WT29.96
Pythia-12b2023.2522.33WT210.54
Pythia-6.9b2023.2522.09WT211.41
Pythia-160m2023.2520.46WT233.43
Pythia-1b2023.2521.26WT216.45
Pythia-1.4b2023.2521.40WT214.72
Pythia-410m2023.2520.87WT220.11
Pythia-2.8b2023.2521.70WT212.69
Sparse Wide GPT-3 Small2023.2219.95WT10320.4
LLaMA-33B2023.1623.48WT26.9
LLaMA-13B2023.1622.93WT213.99
LLaMA-7B2023.1622.66WT29.49
LLaMA-65B2023.1623.78WT24.96
GPT-2+Active-SGD2023.0617.49WT220.59
Hybrid H3-355M2022.9920.05WT10316.9
Hybrid H3-125M2022.9919.59WT10323.7
Hybrid H3-2.7B2022.9920.93WT10310.6
Hybrid H3-1.3B2022.9920.61WT10312.5
Transformer + GFM2022.9118.91WT10320.05
Mogrifier RLSTM (PTB)2022.8416.73PTB42.9
Mogrifier RLSTM (WT2)2022.8417.04WT238
Decaying Fast Weights Transformer2022.7719.11WT10320.5
NMST+GPT-22022.7520.08WT10320.69
BLOOM-1.7B2022.5121.56WT220.17
BLOOM-1B2022.5121.35WT223.7
BLOOM-560M2022.5121.07WT230.05
BLOOM-3B2022.5121.80WT217.57
BLOOM-7.1B2022.5122.17WT214.72
OPT-125M (finetuned on PTB)2022.4720.35PTB16.5
OPT-2.7B (finetuned on PTB)2022.4721.69PTB10.8
OPT-1.3B (finetuned on PTB)2022.4721.37PTB12.02
OPT-66B2022.4723.08WT29.34
OPT-2.7B (finetuned on WT2)2022.4721.69WT210.27
OPT-6.7B2022.4722.08WT210.86
OPT-1.3B (finetuned)2022.4721.37WT212.22
OPT-2.7B2022.4721.69WT212.47
OPT-175B2022.4723.62WT28.35
OPT-13B2022.4722.37WT210.13
OPT-125M (finetuned)2022.4720.35WT219.85
OPT-1.3B2022.4721.37WT216.41
OPT-30B2022.4722.73WT210.67
OPT-350M2022.4720.35WT225.42
DITTO2022.4319.04WT10324.33
B2T connection (16L)2022.4119.45WT10319.2
GPT-NeoX-20B2022.2822.75WT29.2
LaMemo2022.2818.87WT10323.77
Monarch-GPT-2-Medium2022.2520.64WT10320.3
Monarch-GPT-2-Small2022.2520.28WT10320.7
Chinchilla2022.2423.76WT1037.16
NoPos2022.2420.21WT10320.97
Segatron-XL large, M=384 + HCP2022.2219.42WT10317
Transformer Large + HCP2022.2218.78WT10325.3
Segatron -XL base, M=150 + HCP2022.2218.24WT10322.1
MemSizer2022.2218.86WT10320.8
GPT3-6.7B + muP2022.1822.11WT1038.56
HSO2021.9620.54WT10320.3
Gopher (7.1B)2021.9323.80WT10310.81
Gopher (280B)2021.9322.11WT1038.12
GPT-2-Medium+Pixelfly2021.9119.10WT10321
GPT-2-Small+Pixelfly2021.9118.62WT10322.5
GPT2+CoreLM+Fine-Tuning2021.8416.50WT10329.51
GPT2+CoreLM+Fine-Tuning2021.8416.50WT231.8
S42021.8319.89WT10320.95
GPT-2 (fine-tuned with HYDRA)2021.7916.28WT215.17
base LM+GNN+kNN2021.7918.86WT10316.8
PermuteFormer2021.6818.49WT10332.49
$\infty$-former (SM)2021.6720.08WT10316.61
ALiBi (L=3072, Lvalid = 3072)2021.6520.26WT10318.3
GPT-2 (1.5B, Curriculum Learning 45K)2021.6120.78WT10313.72
DEQ-Transformer (Post-LN) + Jacobian Regularisation2021.4919.46WT10324.9
Adaptive Input Transformer + RD2021.4919.91WT10318.07
GPT-J-6B2021.4422.16WT210.88
Delta RNN (+ full context)2021.4418.04WT10332.8
Transformer-C2021.2718.26WT10325.1
GPT-Neo-2.7B (finetuned on PTB)2021.2221.81PTB14.7
GPT-Neo-2.7B (finetuned)2021.2221.81WT210.78
GPT-Neo-125M(finetuned)2021.2220.48WT221.96
GPT-Neo-125M2021.2220.48WT232.29
GPT-Neo-2.7B2021.2221.81WT211.39
GPT-Neo-1.3B (finetuned)2021.2221.81WT212.09
GLM-10B-bidirectional2021.2122.58WT10311.33
GLM-10B-unidirectional2021.2122.58WT10312.22
RFA-GATE-Gaussian-Stateful Big2021.1718.85WT10323.5
SRU++ Large2021.1519.04WT10317.1
SRU++ Large only 2 attention layers (k=5)2021.1518.90WT10317.3
SRU++ Base2021.1518.76WT10318.3
Linear Transformer (large)2021.1418.59WT10331.5
Linear Transformer (small)2021.1418.47WT10335.5
Selfish-RNN (ON-LSTM)2021.0616.80PTB55.82
Selfish-RNN (SNT-ASGD) Stacked LSTMs2021.0616.15PTB71.42
Selfish-RNN (SNT-ASGD)RHNs2021.0616.33PTB64.03
Selfish-RNN (AWD-LSTM-MoS)2021.0617.29WT263.05
Shortformer2021.0018.48WT10318.15
ERNIE-Doc (151M)2021.0019.25WT10321
ERNIE-Doc (247M)2021.0019.46WT10316.8
Subformer (83M)2021.0018.56WT10320.88
Subformer (122M)2021.0018.72WT10319.9
Subformer (96M)2021.0018.62WT10320.39
CT-MoS (PTB)2020.9817.13PTB54.69
CT-MoS + DynamicEval (PTB)2020.9817.13PTB47.42
CT-MoS + DynamicEval (WT2)2020.9817.75WT240.96
CT-MoS (WT2)2020.9817.75WT262.21
AWD-FWM (PTB)2020.8817.13PTB54.48
AWD-FWM (WT2)2020.8817.87WT261.65
Transformer+Recurrent Windows of Context2020.6220.07WT10326.73
DeLight2020.5919.38WT10324.14
3-Layer-Tensor-Transformer+AdaHessian2020.4215.30PTB51.5
6-Layer-Tensor-Transformer+AdaHessian2020.4218.20WT10319.9
GPT3-6.7B (rerun of original)2020.4122.08WT1039.13
rTop-k(distributed setting)2020.3916.16PTB82.49
ONLSTM-SYD2020.3617.14PTB55.7
Segatron XL base, M=3842020.3318.24WT10322.5
Segatron XL large, M=3842020.3319.42WT10317.1
DiffStk-MRNN2020.2614.45PTB115
Tensor-Transformer(1core)+PN (PTB)2020.2115.30PTB47.6
Tensor-Transformer(1core)+PN (WT103)2020.2118.20WT10317.9
TransformerXL + spectrum control2020.1917.66WT10323.2
LSTM-3-layer+Gadam2020.1716.43PTB58.77
Feedback Transformer2020.1419.32WT10318.3
Turing-NLG2020.1222.20WT10310.21
TaLK Convolution2020.1019.44WT10323.3
bRSM + cache2019.9214.33PTB103.5
AWD-LSTM + DeFINE2019.9015.35PTB54.2
Transformer-XL DeFINE (107M)2019.9018.72WT10325.72
Adaptive LSTM + DeFINE2019.9018.79WT10335.94
Transformer-XL DeFINE (141M)2019.9018.79WT10324.17
Compressive Transformers for Long-Range Sequence Modelling2019.8720.20WT10317.1
Sandwich Transformer2019.8620.20WT10317.84
LSTM(medium)+Sememe+cell2019.8015.70WT289.16
Megatron-LM (8.3B)2019.7121.96WT10310.81
Megatron-LM (355M)2019.7120.64WT10319.31
DEQ-TrellisNet2019.6717.91PTB57.1
DEQ-Transformer (Medium, Adaptive Embedding)2019.6717.91WT10323.2
R-Transformer2019.5315.92PTB84.38
All-attention network + adaptive span2019.5019.66WT10320.6
Tensorized Transformer (small)2019.4815.30PTB57.9
Tensorized Transformer (large PTB)2019.4815.60PTB52.7
Tensorized Transformer (257M)2019.4818.68WT10321.2
Tensorized Transformer (core-2)2019.4818.20WT10318.9
Tensorized Transformer (151M)2019.4818.45WT10318.8
Adversarial + AWD-LSTM-MoS + partial shuffled2019.4416.74PTB46.01
AdvSoft + 4 layer QRNN + dynamic evaluation2019.4417.56WT10328
4 layer QRNN + dynamic evaluation2019.4417.56WT10331.6
AWD-LSTM + MoS + Partial Shuffled2019.4417.52WT238.07
Transformer-XL Large + Phrase Induction2019.4217.20WT10317.4
AWD-LSTM-DRILL + dynamic evaluation† (PTB)2019.3617.13PTB49.4
AWD-LSTM-DRILL + dynamic evaluation† (WT2)2019.3617.63WT242
GPT-2 (1542M)2019.1221.18PTB35.76
GPT-2 (1542M)2019.1221.18WT10317.48
GPT-2 (762M)2019.1220.88WT10322.05
GPT-2 (345M)2019.1220.54WT10326.37
GPT-2 (1542M)2019.1221.18WT218.34
Transformer-XL-ptb2019.0219.04PTB54.52
Transformer-XL Large2019.0219.04WT10318.3
Multi-cell LSTM2018.8715.30PTB77.12
Fine-tuned-AWD-LSTM-DOC(fin)2018.8615.28PTB52.12
TrellisNet-MoS (1.4x larger)2018.7918.44PTB54.19
TrellisNet2018.7918.44WT10329.19
TrellisNet-MoS (1.4x larger)2018.7918.44WT10329.19
Transformer (Adaptive Input Embeddings)2018.7418.86WT10318.7
LSTM+NeuralCache2018.7315.01WT266.2
AWD-LSTM-DOC (fin) (23M)2018.6616.94PTB52.38
AWD-LSTM-DOC (fin) (37M)2018.6617.14WT258.03
AWD-LSTM-MoS+PDR + dynamic evaluation (PTB)2018.6217.15PTB47.3
aLSTM(depth-2)+RecurrentPolicy (PTB)2018.3916.38PTB55.3
aLSTM(depth-2)+RecurrentPolicy (WT2)2018.3916.88WT264.5
AWD-LSTM-MoS+Noisin+dynamic evaluation 2018.3316.69PTB47.6
Dropout-LSTM+Noise(Bernoulli) (PTB)2018.3316.76PTB66.1
LSTM+Noise(Beta)2018.3317.10WT282.9
Dropout-LSTM+Noise(Laplace)2018.3316.51WT282.1
Dropout-LSTM+Noise(Bernoulli) (WT2)2018.3317.10WT276.8
LSTM (Hebbian, Cache, MbPA)2018.2319.38WT10329.2
4 layer QRNN (h=2500)2018.2217.38WT10333
QRNN2018.0817.56WT10333
RNNLM + Dynamic KL Regularization2018.0015.49PTB77.8
RNNLM + Dynamic KL Regularization (WT2)2018.0016.34WT286.8
AWD-LSTM-MoS + dynamic evaluation (PTB, 2017)2017.8617.09PTB47.69
AWD-LSTM-MoS + dynamic evaluation (WT2, 2017)2017.8617.64WT240.68
Fraternal dropout + AWD-LSTM 3-layer (PTB)2017.8316.84PTB56.8
Fraternal dropout + AWD-LSTM 3-layer (WT2)2017.8317.34WT264.1
AWD-LSTM+WT+Cache+IOG (PTB)2017.7314.92PTB53
AWD-LSTM+WT+Cache+IOG (WT2)2017.7315.52WT251.7
GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (PTB)2017.6617.16PTB46.34
GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (WT2)2017.6617.68WT240.46
EI-REHN-1200D2017.6215.83PTB66.2
EI-REHN-1000D2017.6216.03PTB68.7
AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (PTB)2017.6016.83PTB52.8
AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (WT2)2017.6017.49WT252
4 layer Densely Connected LSTM2017.5415.89PTB76.8
Densely Connected LSTM + Var. Dropout2017.5416.11PTB78.3
GCRN-M1, dropout2016.9715.48PTB98.67
VD-LSTM+REAL Large2016.8416.33PTB68.5
Pointer Sentinel-LSTM (medium)2016.7415.87PTB70.9
Zoneout + Variational LSTM (PTB)2016.7415.87PTB80.6
Zoneout + Variational LSTM (WT2)2016.7416.23WT2100.9
Pointer Sentinel-LSTM2016.7416.20WT280.8
Variational RHN + WT2016.5315.41PTB65.4
VD-RHN2016.5315.55PTB68.5
Variational (untied weights, MC) LSTM (Large)2015.9615.75PTB73.4
LSTM-Char-Large2015.6515.42PTB78.9
Search-Proven Best LSTM2015.5115.52PTB79.83
genCNN + dyn eval2015.2116.86PTB106.3

Turning one row of that table into terms of the equation

The file gives four things about each result. The equation asks for the same four in a different shape. Nothing is added, estimated or filled in: every line below is one column rewritten. Here is the newest WT103 row doing it.

what the file has
score 20.4
what the equation needs
log10 P
worked out
log10 20.4 = 1.310
what the file has
training compute, 1019.95 FLOP
what the equation needs
log10 C
worked out
the power itself: 19.95
what the file has
dated 2023.22
what the equation needs
gyear
worked out
the whole part: 2023, so this row uses g2023
what the file has
test WT103
what the equation needs
atest
worked out
one of 3 baselines, so this row uses aWT103

Doing that to every row leaves one number per column and nothing else. Not every row gets through, and this is the whole of what is dropped:

fetched from the source212
every published result
with both a compute and a score212
a row missing either cannot be placed on the chart
in years holding 3 results or more208
a year holding one model tells you about that model, not about the year

That last rule drops 4 rows across 2012, 2013, 2014, and it is the one choice on this page that moves the answer. Section 08 says what the number would be without it.

What is left is 208 rows, each carrying a score, a compute, a year and a test. That is one equation per row, and 208 of them to solve at once.

Four steps. Each one: the data, the equation, the result.

Step 1 · finding b

What ten times the compute is worth

The data

The 207 results that share a test and a year with at least one other result, sorted into 22 groups of one test in one year.

Every row in a group was run on the same test in the same year, so atest and gyear are one shared number for all of them. Neither has to be known yet.

The equation

b = Σ x*y / Σ x²b = −33.1370 / 401.0633

x is a row's compute measured from its own group's average, and y is its score measured the same way. Σ means "add up the column", over every row of every group. Both columns are worked out below.

The result

−0.0826

what ten times the compute does to the score, as a power of ten: 10−0.0826 = 0.827, so the score drops by about 17%.

One b, shared by every row and every year. Step 2 takes it back off the data.

Why the rows are grouped first

Rows on the same test in the same year carry the same atest and the same gyear, so inside a group those two are a single constant. Call it k. Every row in the group then reads log10 P = k + b * log10 C, and taking the group's average off both sides removes k entirely. What is left is b, in the last two columns.

So every group has its own b, printed in its heading. They do not agree, and they do not count equally: a group whose models all trained on much the same compute has an x² total near zero and a slope that is mostly noise, so the totals at the foot weight it to nothing. That weight is in the heading too. Years read top to bottom, with the 3 tests side by side inside each one.

Every evaluation, grouped by test and year, with the sums each group contributes
modelits equationx = log C − group avgy = log P − group avgx * yx²
PTB · 20154 rows shared k = aPTB + g2015 avg log C 15.89 avg log P 1.923
b = 0.09740.3%
LSTM-Char-Large1.897 = k + b * 15.42−0.467−0.0260.01210.2186
Search-Proven Best LSTM1.902 = k + b * 15.52−0.367−0.0210.00760.1351
Variational (untied weights, MC) LSTM (Large)1.866 = k + b * 15.75−0.137−0.0570.00790.0189
genCNN + dyn eval2.027 = k + b * 16.860.9730.1040.10080.9458
added up over 4 rows0.12831.3183
b = 0.1283 ÷ 1.3183 = 0.0974 weight = 1.3183 ÷ 401.0633 = 0.0033
PTB · 20166 rows shared k = aPTB + g2016 avg log C 15.75 avg log P 1.873
b = −0.04390.1%
Variational RHN + WT1.816 = k + b * 15.41−0.342−0.0570.01960.1167
GCRN-M1, dropout1.994 = k + b * 15.48−0.2720.121−0.03290.0738
VD-RHN1.836 = k + b * 15.55−0.202−0.0370.00750.0407
Pointer Sentinel-LSTM (medium)1.851 = k + b * 15.870.118−0.022−0.00260.0140
Zoneout + Variational LSTM (PTB)1.906 = k + b * 15.870.1180.0330.00390.0140
VD-LSTM+REAL Large1.836 = k + b * 16.330.578−0.037−0.02160.3345
added up over 6 rows−0.02610.5937
b = −0.0261 ÷ 0.5937 = −0.0439 weight = 0.5937 ÷ 401.0633 = 0.0015
WT2 · 20162 rows shared k = aWT2 + g2016 avg log C 16.21 avg log P 1.956
b = 3.21600.0%
Pointer Sentinel-LSTM1.907 = k + b * 16.20−0.015−0.0480.00070.0002
Zoneout + Variational LSTM (WT2)2.004 = k + b * 16.230.0150.0480.00070.0002
added up over 2 rows0.00140.0005
b = 0.0014 ÷ 0.0005 = 3.2160 weight = 0.0005 ÷ 401.0633 = 0.0000
PTB · 20179 rows shared k = aPTB + g2017 avg log C 16.30 avg log P 1.776
b = −0.05651.1%
AWD-LSTM+WT+Cache+IOG (PTB)1.724 = k + b * 14.92−1.380−0.0520.07121.9044
EI-REHN-1200D1.821 = k + b * 15.83−0.4700.045−0.02120.2209
4 layer Densely Connected LSTM1.885 = k + b * 15.89−0.4100.110−0.04490.1681
EI-REHN-1000D1.837 = k + b * 16.03−0.2700.061−0.01650.0729
Densely Connected LSTM + Var. Dropout1.894 = k + b * 16.11−0.1900.118−0.02240.0361
AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (PTB)1.723 = k + b * 16.830.530−0.053−0.02820.2809
Fraternal dropout + AWD-LSTM 3-layer (PTB)1.754 = k + b * 16.840.540−0.021−0.01160.2916
AWD-LSTM-MoS + dynamic evaluation (PTB, 2017)1.678 = k + b * 17.090.790−0.097−0.07700.6241
GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (PTB)1.666 = k + b * 17.160.860−0.110−0.09450.7396
added up over 9 rows−0.24514.3386
b = −0.2451 ÷ 4.3386 = −0.0565 weight = 4.3386 ÷ 401.0633 = 0.0108
WT2 · 20175 rows shared k = aWT2 + g2017 avg log C 17.13 avg log P 1.691
b = −0.02720.8%
AWD-LSTM+WT+Cache+IOG (WT2)1.713 = k + b * 15.52−1.6140.023−0.03702.6050
Fraternal dropout + AWD-LSTM 3-layer (WT2)1.807 = k + b * 17.340.2060.1160.02400.0424
AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (WT2)1.716 = k + b * 17.490.3560.0250.00910.1267
AWD-LSTM-MoS + dynamic evaluation (WT2, 2017)1.609 = k + b * 17.640.506−0.081−0.04110.2560
GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (WT2)1.607 = k + b * 17.680.546−0.084−0.04560.2981
added up over 5 rows−0.09073.3283
b = −0.0907 ÷ 3.3283 = −0.0272 weight = 3.3283 ÷ 401.0633 = 0.0083
PTB · 20189 rows shared k = aPTB + g2018 avg log C 16.49 avg log P 1.763
b = −0.04192.1%
Fine-tuned-AWD-LSTM-DOC(fin)1.717 = k + b * 15.28−1.212−0.0460.05531.4695
Multi-cell LSTM1.887 = k + b * 15.30−1.1920.125−0.14851.4214
RNNLM + Dynamic KL Regularization1.891 = k + b * 15.49−1.0020.128−0.12861.0044
aLSTM(depth-2)+RecurrentPolicy (PTB)1.743 = k + b * 16.38−0.112−0.0200.00220.0126
AWD-LSTM-MoS+Noisin+dynamic evaluation 1.678 = k + b * 16.690.198−0.085−0.01680.0391
Dropout-LSTM+Noise(Bernoulli) (PTB)1.820 = k + b * 16.760.2680.0580.01540.0717
AWD-LSTM-DOC (fin) (23M)1.719 = k + b * 16.940.448−0.043−0.01950.2005
AWD-LSTM-MoS+PDR + dynamic evaluation (PTB)1.675 = k + b * 17.150.658−0.088−0.05770.4327
TrellisNet-MoS (1.4x larger)1.734 = k + b * 18.441.948−0.029−0.05593.7938
added up over 9 rows−0.35418.4458
b = −0.3541 ÷ 8.4458 = −0.0419 weight = 8.4458 ÷ 401.0633 = 0.0211
WT103 · 20186 rows shared k = aWT103 + g2018 avg log C 18.34 avg log P 1.451
b = −0.06640.7%
4 layer QRNN (h=2500)1.519 = k + b * 17.38−0.9630.068−0.06520.9280
QRNN1.519 = k + b * 17.56−0.7830.068−0.05310.6136
TrellisNet1.465 = k + b * 18.440.0970.0140.00140.0093
TrellisNet-MoS (1.4x larger)1.465 = k + b * 18.440.0970.0140.00140.0093
Transformer (Adaptive Input Embeddings)1.272 = k + b * 18.860.517−0.179−0.09250.2669
LSTM (Hebbian, Cache, MbPA)1.465 = k + b * 19.381.0370.0150.01511.0747
added up over 6 rows−0.19282.9019
b = −0.1928 ÷ 2.9019 = −0.0664 weight = 2.9019 ÷ 401.0633 = 0.0072
WT2 · 20187 rows shared k = aWT2 + g2018 avg log C 16.58 avg log P 1.864
b = 0.00380.9%
LSTM+NeuralCache1.821 = k + b * 15.01−1.573−0.0440.06852.4739
RNNLM + Dynamic KL Regularization (WT2)1.939 = k + b * 16.34−0.2430.074−0.01800.0590
Dropout-LSTM+Noise(Laplace)1.914 = k + b * 16.51−0.0730.050−0.00360.0053
aLSTM(depth-2)+RecurrentPolicy (WT2)1.810 = k + b * 16.880.297−0.055−0.01630.0883
LSTM+Noise(Beta)1.919 = k + b * 17.100.5170.0540.02800.2674
Dropout-LSTM+Noise(Bernoulli) (WT2)1.885 = k + b * 17.100.5170.0210.01080.2674
AWD-LSTM-DOC (fin) (37M)1.764 = k + b * 17.140.557−0.101−0.05610.3104
added up over 7 rows0.01333.4717
b = 0.0133 ÷ 3.4717 = 0.0038 weight = 3.4717 ÷ 401.0633 = 0.0087
PTB · 201910 rows shared k = aPTB + g2019 avg log C 16.85 avg log P 1.756
b = −0.04389.5%
bRSM + cache2.015 = k + b * 14.33−2.5200.259−0.65186.3504
Tensorized Transformer (small)1.763 = k + b * 15.30−1.5500.006−0.00992.4025
AWD-LSTM + DeFINE1.734 = k + b * 15.35−1.500−0.0220.03342.2500
Tensorized Transformer (large PTB)1.722 = k + b * 15.60−1.250−0.0340.04311.5625
R-Transformer1.926 = k + b * 15.92−0.9300.170−0.15810.8649
Adversarial + AWD-LSTM-MoS + partial shuffled1.663 = k + b * 16.74−0.110−0.0930.01030.0121
AWD-LSTM-DRILL + dynamic evaluation† (PTB)1.694 = k + b * 17.130.280−0.063−0.01750.0784
DEQ-TrellisNet1.757 = k + b * 17.911.0600.0000.00041.1236
Transformer-XL-ptb1.737 = k + b * 19.042.190−0.020−0.04324.7961
GPT-2 (1542M)1.553 = k + b * 21.184.330−0.203−0.878518.7489
added up over 10 rows−1.671838.1894
b = −1.6718 ÷ 38.1894 = −0.0438 weight = 38.1894 ÷ 401.0633 = 0.0952
WT103 · 201919 rows shared k = aWT103 + g2019 avg log C 19.27 avg log P 1.325
b = −0.04348.4%
Transformer-XL Large + Phrase Induction1.241 = k + b * 17.20−2.072−0.0840.17404.2914
AdvSoft + 4 layer QRNN + dynamic evaluation1.447 = k + b * 17.56−1.7120.123−0.20992.9295
4 layer QRNN + dynamic evaluation1.500 = k + b * 17.56−1.7120.175−0.29982.9295
DEQ-Transformer (Medium, Adaptive Embedding)1.365 = k + b * 17.91−1.3620.041−0.05571.8539
Tensorized Transformer (core-2)1.276 = k + b * 18.20−1.072−0.0480.05151.1483
Tensorized Transformer (151M)1.274 = k + b * 18.45−0.822−0.0500.04140.6750
Tensorized Transformer (257M)1.326 = k + b * 18.68−0.5920.002−0.00110.3500
Transformer-XL DeFINE (107M)1.410 = k + b * 18.72−0.5520.086−0.04730.3042
Adaptive LSTM + DeFINE1.556 = k + b * 18.79−0.4820.231−0.11130.2319
Transformer-XL DeFINE (141M)1.383 = k + b * 18.79−0.4820.059−0.02830.2319
Transformer-XL Large1.262 = k + b * 19.04−0.232−0.0620.01440.0536
All-attention network + adaptive span1.314 = k + b * 19.660.388−0.011−0.00410.1509
Sandwich Transformer1.251 = k + b * 20.200.928−0.073−0.06790.8620
Compressive Transformers for Long-Range Sequence Modelling1.233 = k + b * 20.200.928−0.092−0.08500.8620
GPT-2 (345M)1.421 = k + b * 20.541.2680.0970.12251.6089
Megatron-LM (355M)1.286 = k + b * 20.641.368−0.039−0.05301.8726
GPT-2 (762M)1.343 = k + b * 20.881.6080.0190.03032.5870
GPT-2 (1542M)1.243 = k + b * 21.181.908−0.082−0.15653.6421
Megatron-LM (8.3B)1.034 = k + b * 21.962.688−0.291−0.78167.2276
added up over 19 rows−1.467333.8123
b = −1.4673 ÷ 33.8123 = −0.0434 weight = 33.8123 ÷ 401.0633 = 0.0843
WT2 · 20194 rows shared k = aWT2 + g2019 avg log C 18.01 avg log P 1.604
b = −0.11893.9%
LSTM(medium)+Sememe+cell1.950 = k + b * 15.70−2.3080.346−0.79805.3246
AWD-LSTM + MoS + Partial Shuffled1.581 = k + b * 17.52−0.488−0.0240.01160.2377
AWD-LSTM-DRILL + dynamic evaluation† (WT2)1.623 = k + b * 17.63−0.3780.019−0.00710.1425
GPT-2 (1542M)1.263 = k + b * 21.183.172−0.341−1.081710.0648
added up over 4 rows−1.875215.7695
b = −1.8752 ÷ 15.7695 = −0.1189 weight = 15.7695 ÷ 401.0633 = 0.0393
PTB · 20209 rows shared k = aPTB + g2020 avg log C 16.24 avg log P 1.781
b = −0.06772.0%
DiffStk-MRNN2.061 = k + b * 14.45−1.7910.279−0.50043.2081
Tensor-Transformer(1core)+PN (PTB)1.678 = k + b * 15.30−0.941−0.1040.09760.8857
3-Layer-Tensor-Transformer+AdaHessian1.712 = k + b * 15.30−0.941−0.0690.06540.8857
rTop-k(distributed setting)1.916 = k + b * 16.16−0.0810.135−0.01100.0066
LSTM-3-layer+Gadam1.769 = k + b * 16.430.189−0.012−0.00230.0357
AWD-FWM (PTB)1.736 = k + b * 17.130.889−0.045−0.04000.7901
CT-MoS (PTB)1.738 = k + b * 17.130.889−0.043−0.03860.7901
CT-MoS + DynamicEval (PTB)1.676 = k + b * 17.130.889−0.105−0.09360.7901
ONLSTM-SYD1.746 = k + b * 17.140.899−0.035−0.03190.8080
added up over 9 rows−0.55488.2001
b = −0.5548 ÷ 8.2001 = −0.0677 weight = 8.2001 ÷ 401.0633 = 0.0204
WT103 · 202014 rows shared k = aWT103 + g2020 avg log C 19.39 avg log P 1.266
b = −0.07335.9%
TransformerXL + spectrum control1.365 = k + b * 17.66−1.7260.100−0.17242.9781
Tensor-Transformer(1core)+PN (WT103)1.253 = k + b * 18.20−1.186−0.0130.01511.4059
6-Layer-Tensor-Transformer+AdaHessian1.299 = k + b * 18.20−1.1860.033−0.03951.4059
Segatron XL base, M=3841.352 = k + b * 18.24−1.1460.087−0.09921.3127
Shortformer1.259 = k + b * 18.48−0.906−0.0070.00610.8203
ERNIE-Doc (151M)1.322 = k + b * 19.25−0.1360.057−0.00770.0184
Feedback Transformer1.262 = k + b * 19.32−0.066−0.0030.00020.0043
DeLight1.383 = k + b * 19.38−0.0060.117−0.00070.0000
Segatron XL large, M=3841.233 = k + b * 19.420.034−0.033−0.00110.0012
TaLK Convolution1.367 = k + b * 19.440.0540.1020.00550.0029
ERNIE-Doc (247M)1.225 = k + b * 19.460.074−0.040−0.00300.0055
Transformer+Recurrent Windows of Context1.427 = k + b * 20.070.6840.1610.11050.4682
GPT3-6.7B (rerun of original)0.960 = k + b * 22.082.694−0.305−0.82207.2592
Turing-NLG1.009 = k + b * 22.202.814−0.257−0.72207.9202
added up over 14 rows−1.730323.6029
b = −1.7303 ÷ 23.6029 = −0.0733 weight = 23.6029 ÷ 401.0633 = 0.0589
WT2 · 20203 rows shared k = aWT2 + g2020 avg log C 17.79 avg log P 1.732
b = 0.72350.0%
CT-MoS + DynamicEval (WT2)1.612 = k + b * 17.75−0.040−0.1200.00480.0016
CT-MoS (WT2)1.794 = k + b * 17.75−0.0400.062−0.00250.0016
AWD-FWM (WT2)1.790 = k + b * 17.870.0800.0580.00460.0064
added up over 3 rows0.00690.0096
b = 0.0069 ÷ 0.0096 = 0.7235 weight = 0.0096 ÷ 401.0633 = 0.0000
PTB · 20214 rows shared k = aPTB + g2021 avg log C 17.77 avg log P 1.644
b = −0.11845.5%
Selfish-RNN (SNT-ASGD) Stacked LSTMs1.854 = k + b * 16.15−1.6230.210−0.34112.6325
Selfish-RNN (SNT-ASGD)RHNs1.806 = k + b * 16.33−1.4430.163−0.23482.0808
Selfish-RNN (ON-LSTM)1.747 = k + b * 16.80−0.9730.103−0.10040.9458
GPT-Neo-2.7B (finetuned on PTB)1.167 = k + b * 21.814.037−0.476−1.922916.3014
added up over 4 rows−2.599221.9605
b = −2.5992 ÷ 21.9605 = −0.1184 weight = 21.9605 ÷ 401.0633 = 0.0548
WT103 · 202127 rows shared k = aWT103 + g2021 avg log C 19.57 avg log P 1.291
b = −0.078917.2%
GPT2+CoreLM+Fine-Tuning1.470 = k + b * 16.50−3.0690.179−0.54849.4204
Delta RNN (+ full context)1.516 = k + b * 18.04−1.5290.225−0.34352.3386
Transformer-C1.400 = k + b * 18.26−1.3090.108−0.14191.7142
Linear Transformer (small)1.550 = k + b * 18.47−1.0990.259−0.28461.2084
PermuteFormer1.512 = k + b * 18.49−1.0790.220−0.23791.1648
Subformer (83M)1.320 = k + b * 18.56−1.0090.028−0.02871.0186
Linear Transformer (large)1.498 = k + b * 18.59−0.9790.207−0.20270.9589
Subformer (96M)1.309 = k + b * 18.62−0.9490.018−0.01720.9011
GPT-2-Small+Pixelfly1.352 = k + b * 18.62−0.9490.061−0.05780.9011
Subformer (122M)1.299 = k + b * 18.72−0.8490.008−0.00640.7212
SRU++ Base1.262 = k + b * 18.76−0.809−0.0290.02330.6549
RFA-GATE-Gaussian-Stateful Big1.371 = k + b * 18.85−0.7190.080−0.05740.5173
base LM+GNN+kNN1.225 = k + b * 18.86−0.709−0.0660.04680.5030
SRU++ Large only 2 attention layers (k=5)1.238 = k + b * 18.90−0.669−0.0530.03560.4479
SRU++ Large1.233 = k + b * 19.04−0.529−0.0580.03090.2801
GPT-2-Medium+Pixelfly1.322 = k + b * 19.10−0.4690.031−0.01450.2202
DEQ-Transformer (Post-LN) + Jacobian Regularisation1.396 = k + b * 19.46−0.1090.105−0.01150.0119
S41.321 = k + b * 19.890.3210.0300.00960.1029
Adaptive Input Transformer + RD1.257 = k + b * 19.910.341−0.034−0.01170.1161
$\infty$-former (SM)1.220 = k + b * 20.080.511−0.071−0.03620.2609
ALiBi (L=3072, Lvalid = 3072)1.262 = k + b * 20.260.691−0.029−0.01990.4771
HSO1.307 = k + b * 20.540.9710.0160.01570.9423
GPT-2 (1.5B, Curriculum Learning 45K)1.137 = k + b * 20.781.211−0.154−0.18641.4659
Gopher (280B)0.910 = k + b * 22.112.541−0.382−0.96996.4554
GLM-10B-bidirectional1.054 = k + b * 22.583.011−0.237−0.71379.0646
GLM-10B-unidirectional1.087 = k + b * 22.583.011−0.204−0.61489.0646
Gopher (7.1B)1.034 = k + b * 23.804.231−0.257−1.089317.8992
added up over 27 rows−5.432668.8316
b = −5.4326 ÷ 68.8316 = −0.0789 weight = 68.8316 ÷ 401.0633 = 0.1716
WT2 · 20219 rows shared k = aWT2 + g2021 avg log C 19.85 avg log P 1.282
b = −0.070812.0%
GPT-2 (fine-tuned with HYDRA)1.181 = k + b * 16.28−3.567−0.1010.361912.7211
GPT2+CoreLM+Fine-Tuning1.502 = k + b * 16.50−3.3470.220−0.736211.2002
Selfish-RNN (AWD-LSTM-MoS)1.800 = k + b * 17.29−2.5570.517−1.32246.5365
GPT-Neo-125M(finetuned)1.342 = k + b * 20.480.6330.0590.03750.4011
GPT-Neo-125M1.509 = k + b * 20.480.6330.2270.14350.4011
GPT-Neo-2.7B (finetuned)1.033 = k + b * 21.811.963−0.250−0.49053.8547
GPT-Neo-2.7B1.057 = k + b * 21.811.963−0.226−0.44363.8547
GPT-Neo-1.3B (finetuned)1.082 = k + b * 21.811.963−0.200−0.39273.8547
GPT-J-6B1.037 = k + b * 22.162.313−0.246−0.56875.3515
added up over 9 rows−3.411148.1756
b = −3.4111 ÷ 48.1756 = −0.0708 weight = 48.1756 ÷ 401.0633 = 0.1201
PTB · 20224 rows shared k = aPTB + g2022 avg log C 20.04 avg log P 1.241
b = −0.11963.9%
Mogrifier RLSTM (PTB)1.632 = k + b * 16.73−3.3050.392−1.294410.9230
OPT-125M (finetuned on PTB)1.217 = k + b * 20.350.315−0.023−0.00740.0992
OPT-1.3B (finetuned on PTB)1.080 = k + b * 21.371.335−0.161−0.21481.7822
OPT-2.7B (finetuned on PTB)1.033 = k + b * 21.691.655−0.207−0.34322.7390
added up over 4 rows−1.859815.5435
b = −1.8598 ÷ 15.5435 = −0.1196 weight = 15.5435 ÷ 401.0633 = 0.0388
WT103 · 202219 rows shared k = aWT103 + g2022 avg log C 19.94 avg log P 1.249
b = −0.10457.8%
Segatron -XL base, M=150 + HCP1.344 = k + b * 18.24−1.7040.096−0.16282.9043
Transformer Large + HCP1.403 = k + b * 18.78−1.1640.154−0.17961.3554
MemSizer1.318 = k + b * 18.86−1.0840.069−0.07501.1755
LaMemo1.376 = k + b * 18.87−1.0740.127−0.13661.1539
Transformer + GFM1.302 = k + b * 18.91−1.0340.053−0.05511.0696
DITTO1.386 = k + b * 19.04−0.9040.137−0.12410.8176
Decaying Fast Weights Transformer1.312 = k + b * 19.11−0.8340.063−0.05250.6959
Segatron-XL large, M=384 + HCP1.230 = k + b * 19.42−0.524−0.0180.00970.2748
B2T connection (16L)1.283 = k + b * 19.45−0.4940.034−0.01700.2442
Hybrid H3-125M1.375 = k + b * 19.59−0.3540.126−0.04460.1255
Hybrid H3-355M1.228 = k + b * 20.050.106−0.021−0.00220.0112
NMST+GPT-21.316 = k + b * 20.080.1360.0670.00910.0184
NoPos1.322 = k + b * 20.210.2660.0730.01930.0706
Monarch-GPT-2-Small1.316 = k + b * 20.280.3360.0670.02250.1128
Hybrid H3-1.3B1.097 = k + b * 20.610.666−0.152−0.10120.4433
Monarch-GPT-2-Medium1.307 = k + b * 20.640.6960.0590.04080.4841
Hybrid H3-2.7B1.025 = k + b * 20.930.986−0.224−0.22040.9718
GPT3-6.7B + muP0.932 = k + b * 22.112.166−0.316−0.68524.6906
Chinchilla0.855 = k + b * 23.763.816−0.394−1.503214.5602
added up over 19 rows−3.258131.1799
b = −3.2581 ÷ 31.1799 = −0.1045 weight = 31.1799 ÷ 401.0633 = 0.0777
WT2 · 202218 rows shared k = aWT2 + g2022 avg log C 21.58 avg log P 1.177
b = −0.11398.6%
Mogrifier RLSTM (WT2)1.580 = k + b * 17.04−4.5400.403−1.828220.6116
OPT-125M (finetuned)1.298 = k + b * 20.35−1.2300.121−0.14841.5129
OPT-350M1.405 = k + b * 20.35−1.2300.228−0.28051.5129
BLOOM-560M1.478 = k + b * 21.07−0.5100.301−0.15340.2601
BLOOM-1B1.375 = k + b * 21.35−0.2300.198−0.04550.0529
OPT-1.3B (finetuned)1.087 = k + b * 21.37−0.210−0.0900.01890.0441
OPT-1.3B1.215 = k + b * 21.37−0.2100.038−0.00800.0441
BLOOM-1.7B1.305 = k + b * 21.56−0.0200.128−0.00260.0004
OPT-2.7B (finetuned on WT2)1.012 = k + b * 21.690.110−0.166−0.01820.0121
OPT-2.7B1.096 = k + b * 21.690.110−0.081−0.00890.0121
BLOOM-3B1.245 = k + b * 21.800.2200.0680.01490.0484
OPT-6.7B1.036 = k + b * 22.080.500−0.141−0.07060.2500
BLOOM-7.1B1.168 = k + b * 22.170.590−0.009−0.00540.3481
OPT-13B1.006 = k + b * 22.370.790−0.171−0.13550.6241
OPT-30B1.028 = k + b * 22.731.150−0.149−0.17131.3225
GPT-NeoX-20B0.964 = k + b * 22.751.170−0.213−0.24961.3689
OPT-66B0.970 = k + b * 23.081.500−0.207−0.31012.2500
OPT-175B0.922 = k + b * 23.622.040−0.255−0.52104.1616
added up over 18 rows−3.923434.4368
b = −3.9234 ÷ 34.4368 = −0.1139 weight = 34.4368 ÷ 401.0633 = 0.0859
PTB · 20233 rows shared k = aPTB + g2023 avg log C 23.02 avg log P 0.936
b = −0.11870.1%
LLaMA-7B (LoRA finetuned)0.986 = k + b * 22.66−0.3630.050−0.01830.1320
LLaMA-13B (LoRA finetuned)0.937 = k + b * 22.93−0.0930.000−0.00000.0087
LLaMA-33B (LoRA finetuned)0.885 = k + b * 23.480.457−0.051−0.02320.2085
added up over 3 rows−0.04150.3493
b = −0.0415 ÷ 0.3493 = −0.1187 weight = 0.3493 ÷ 401.0633 = 0.0009
WT2 · 202316 rows shared k = aWT2 + g2023 avg log C 22.03 avg log P 1.033
b = −0.12449.1%
GPT-2+Active-SGD1.314 = k + b * 17.49−4.5380.281−1.272920.5889
Pythia-160m1.524 = k + b * 20.46−1.5670.491−0.76962.4571
Pythia-410m1.303 = k + b * 20.87−1.1570.270−0.31281.3398
Pythia-1b1.216 = k + b * 21.26−0.7670.183−0.14050.5891
Pythia-1.4b1.168 = k + b * 21.40−0.6280.135−0.08460.3938
Pythia-2.8b1.103 = k + b * 21.70−0.3280.070−0.02300.1073
Pythia-6.9b1.057 = k + b * 22.090.0630.0240.00150.0039
Pythia-12b1.023 = k + b * 22.330.302−0.010−0.00310.0915
MPT-7B0.998 = k + b * 22.620.593−0.035−0.02070.3511
LLaMA-7B0.977 = k + b * 22.660.633−0.056−0.03530.4001
LLaMA-7B (LoRA finetuned)0.792 = k + b * 22.660.633−0.241−0.15270.4001
LLaMA-13B1.146 = k + b * 22.930.9020.1130.10170.8145
LLaMA-13B (LoRA finetuned)0.744 = k + b * 22.930.902−0.290−0.26140.8145
LLaMA-33B0.839 = k + b * 23.481.453−0.194−0.28222.1098
LLaMA-65B0.695 = k + b * 23.781.753−0.338−0.59173.0713
LLaMA-65B (LoRA finetuned)0.630 = k + b * 23.781.753−0.403−0.70573.0713
added up over 16 rows−4.553136.6037
b = −4.5531 ÷ 36.6037 = −0.1244 weight = 36.6037 ÷ 401.0633 = 0.0913

That b column, drawn against the year:

−0.10−0.050.00201520162017201820192020202120222023all rows together −0.0826PTB 2015 · 4 rows · b = 0.0974 · 0.3% of the weight · off the top of the scalePTB 2016 · 6 rows · b = −0.0439 · 0.1% of the weightPTB 2017 · 9 rows · b = −0.0565 · 1.1% of the weightPTB 2018 · 9 rows · b = −0.0419 · 2.1% of the weightPTB 2019 · 10 rows · b = −0.0438 · 9.5% of the weightPTB 2020 · 9 rows · b = −0.0677 · 2.0% of the weightPTB 2021 · 4 rows · b = −0.1184 · 5.5% of the weightPTB 2022 · 4 rows · b = −0.1196 · 3.9% of the weightPTB 2023 · 3 rows · b = −0.1187 · 0.1% of the weightWT103 2018 · 6 rows · b = −0.0664 · 0.7% of the weightWT103 2019 · 19 rows · b = −0.0434 · 8.4% of the weightWT103 2020 · 14 rows · b = −0.0733 · 5.9% of the weightWT103 2021 · 27 rows · b = −0.0789 · 17.2% of the weightWT103 2022 · 19 rows · b = −0.1045 · 7.8% of the weightWT2 2016 · 2 rows · b = 3.2160 · 0.0% of the weight · off the top of the scaleWT2 2017 · 5 rows · b = −0.0272 · 0.8% of the weightWT2 2018 · 7 rows · b = 0.0038 · 0.9% of the weightWT2 2019 · 4 rows · b = −0.1189 · 3.9% of the weightWT2 2020 · 3 rows · b = 0.7235 · 0.0% of the weight · off the top of the scaleWT2 2021 · 9 rows · b = −0.0708 · 12.0% of the weightWT2 2022 · 18 rows · b = −0.1139 · 8.6% of the weightWT2 2023 · 16 rows · b = −0.1244 · 9.1% of the weight
PTBWT103WT2bigger dot = more of the weight3 off the scale, drawn as hollow triangles: their models all trained on much the same compute, and between them they carry 0.3% of the weight

One b out of those 22

Each group gets a say in proportion to its weight, and a weight is one division: that group's x² total, carried down from the table above, divided by every group's x² added together. That total is 401.0633, in the foot of this table. Then multiply every group's b by its weight and add that column up.

weight = group's x² ÷ 401.0633

x² is how far apart on compute a group's models are: square each row's x from the table above and add them up. A group whose models all trained on much the same compute has almost none of it, so almost none of the say.

Each group's slope, its weight, and the two multiplied together
grouprowsits bits x² ÷ 401.0633 = weight b * weight
PTB · 201540.09741.31830.00330.00032
PTB · 20166−0.04390.59370.0015−0.00006
WT2 · 201623.21600.00050.00000.00000
PTB · 20179−0.05654.33860.0108−0.00061
WT2 · 20175−0.02723.32830.0083−0.00023
PTB · 20189−0.04198.44580.0211−0.00088
WT103 · 20186−0.06642.90190.0072−0.00048
WT2 · 201870.00383.47170.00870.00003
PTB · 201910−0.043838.18940.0952−0.00417
WT103 · 201919−0.043433.81230.0843−0.00366
WT2 · 20194−0.118915.76950.0393−0.00468
PTB · 20209−0.06778.20010.0204−0.00138
WT103 · 202014−0.073323.60290.0589−0.00431
WT2 · 202030.72350.00960.00000.00002
PTB · 20214−0.118421.96050.0548−0.00648
WT103 · 202127−0.078968.83160.1716−0.01355
WT2 · 20219−0.070848.17560.1201−0.00851
PTB · 20224−0.119615.54350.0388−0.00464
WT103 · 202219−0.104531.17990.0777−0.00812
WT2 · 202218−0.113934.43680.0859−0.00978
PTB · 20233−0.11870.34930.0009−0.00010
WT2 · 202316−0.124436.60370.0913−0.01135
added up401.06331.0000−0.0826
b =

−0.0826

what ten times the compute does to the score, from these 207 rows

There is a shorter road to the same number, and it is the one the pipeline drives. Add every group's x * y together, add every group's x² together, and divide once:

b = −33.1370 ÷ 401.0633 = −0.0826

Weighting each group's b by its x² is the same arithmetic as pooling those two totals, written out one group at a time. That is why the column above adds to exactly this figure rather than merely close to it, and it is the check that the long way and the short way are one calculation.

That is the b the rest of the page uses, and it is the one the model uses too. Every comparison behind it is between two rows on the same test in the same year, which is the only comparison that needs nothing assumed about how a test's baseline and a year's gain combine.

Step 2 · finding atest and gyear

What is left in a score once the compute is off it

The data

All 208 rows again, each with the compute term from step 1 taken off it: log10 P − b * log10 C. Call that a row's leftover.

A leftover can only be two things, because they are the only two left in the equation: how hard that test is, and what that year had learned.

The equation

atest = that test's leftovers, each less its own year's gyear, averaged

gyear = that year's leftovers, each less its own test's atest, averaged

Two plain averages, and each one needs the other's answer first. So start with every gyear at zero and run them in turn until neither moves: 8 passes here.

The result

−0.3536

the 2023 gain: on the same test, for the same compute, a 2023 score sits that far below a 2015 one, in powers of ten.

9 of these, one a year, and one baseline for each of the 3 tests. 2015 reads zero because every other year is measured against it.

log10 P − b * log10 C = atest + gyear

Take the compute term off a row and what is left can only be those two. So take that year's gain off as well, and a test's baseline is simply the average of its own rows.

Neither of those two is known yet, and each one needs the other: a test's baseline needs that year's gain taken off first, and a year's gain needs that test's baseline taken off first. Nothing has to go first, though. Start by assuming there was no progress at all, every gyear at zero, and run the two averages in turn.

Pass 1, first half

With every gain at zero there is nothing to take off, so a test's baseline is the plain average of log10 P − b * log10 C over its own rows.

PTB

3.2355

WT103

3.0314

WT2

3.1254

Pass 1, second half

Take those three off instead, and average within each year. That is a first set of gains, and they are all too small: 2023 comes out at −0.2811 because the baselines it used were built on the assumption that no year had gained anything. So go round again with these gains in hand.

Each pass moves the column less than the one before it, and by pass 6 none of these four decimals changes again. The loop runs 8 passes because it keeps going until a pass moves nothing anywhere by more than a hundred-thousandth, which is finer than what is printed here. The bottom row is what the model actually runs: one least-squares solve that jumps straight to the place this loop walks to. The two agree to within 0.000001, and the three baselines settle on the same values too, which is why the loop is worth showing and the solve is worth using.

year

2023, pass by pass

Every number in that table is one average over real rows, the solve row included. Pick any cell and the rows behind it open below.

Pass 1 · 2023

−0.2811

one average over the 20 results dated 2023, then shifted so 2015 reads zero.

A · the baselines this cell was measured against

This pass started with every gain at zero, so there was nothing to take off: each baseline is the plain average of that test's leftovers.

the gains it started from

2015 0.00002016 0.00002017 0.00002018 0.00002019 0.00002020 0.00002021 0.00002022 0.00002023 0.0000
Each test's baseline for this pass, as the division that makes it
testrows added up, each row's leftover , nothing taken off÷ rows= baseline
58180.361583.1097
86249.875862.9055
64191.970642.9995

These are the baselines step B measures against. The three cards higher up print the same ones with this pass's shift already added, so PTB reads 3.2355 there and 3.1097 here. The gap between the two is step C, and taking it off here or off the year below gives the same cell.

Each baseline is a sum over that test's own rows. Open one and they are listed below.

B · every 2023 row, against its own test's baseline

These are the 20 results the cell is an average of. Each one's leftover is its score as a power of ten with the compute term taken off, and the last column takes off the baseline of the test it was run on. Open a row and it shows its own six pieces of arithmetic, starting from the two numbers the file holds for it.

Every 2023 result, its leftover, and the baseline taken off it
modeltestlog Pb * log Cleftoverits baselineleftover − baseline
WT21.3137−1.44512.75872.9995−0.2408
WT20.8388−1.94002.77882.9995−0.2207
WT21.1458−1.89453.04042.99950.0408
WT20.9773−1.87222.84952.9995−0.1500
WT20.6955−1.96482.66032.9995−0.3393
WT1031.3096−1.64832.95802.90550.0524
WT21.0228−1.84502.86782.9995−0.1317
WT21.0573−1.82512.88242.9995−0.1171
WT21.5241−1.69053.21462.99950.2151
WT21.2162−1.75662.97272.9995−0.0268
WT21.1679−1.76812.93602.9995−0.0635
WT21.3034−1.72433.02782.99950.0282
WT21.1035−1.79292.89642.9995−0.1031
WT20.9983−1.86892.86722.9995−0.1323
PTB0.8854−1.94002.82533.1097−0.2843
PTB0.9365−1.89452.83113.1097−0.2786
PTB0.9863−1.87222.85863.1097−0.2511
WT20.7435−1.89452.63812.9995−0.3615
WT20.6304−1.96482.59522.9995−0.4043
WT20.7917−1.87222.66392.9995−0.3356
added up over 20 rows
−3.104
÷ 20 rows
−0.1552

The total is the pipeline's own, added before anything was rounded, so adding the printed column by hand lands within a thousandth of it.

C · pin 2015 at zero

−0.1552 − 0.1259 = −0.2811

The number taken off is what step B gives for 2015 on this same pass. Taking it off every year is what makes the 2015 column read zero and every other cell a reading against it. It moves the whole row together, so no gap between two years changes.

Every row of a given year carries that year's settled number, which is why the gyear column below repeats down each year while the other two columns change row by row. g2015 reads zero in every pass because each pass is shifted to put it there. The whole column can slide up or down together without changing a single gap between two years, so one year has to be pinned before the numbers mean anything, and pinning 2015 is what makes every gain a reading against it. The division below then arrives at that zero on its own, which is the subject of the last paragraph in this step.

Every row's leftover, filed under the test whose baseline it feeds
modelyearlog Pb * log Cgyearleftover
PTB · 58 rows
genCNN + dyn eval20152.0265−1.39300.00003.4196
Search-Proven Best LSTM20151.9022−1.28230.00003.1845
LSTM-Char-Large20151.8971−1.27400.00003.1711
Variational (untied weights, MC) LSTM (Large)20151.8657−1.30130.00003.1670
Variational RHN + WT20161.8156−1.2732−0.02523.1140
VD-RHN20161.8357−1.2848−0.02523.1457
Pointer Sentinel-LSTM (medium)20161.8506−1.3112−0.02523.1871
Zoneout + Variational LSTM (PTB)20161.9063−1.3112−0.02523.2428
VD-LSTM+REAL Large20161.8357−1.3492−0.02523.2101
GCRN-M1, dropout20161.9942−1.2790−0.02523.2984
4 layer Densely Connected LSTM20171.8854−1.3129−0.11073.3090
Densely Connected LSTM + Var. Dropout20171.8938−1.3311−0.11073.3355
AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (PTB)20171.7226−1.3905−0.11073.2239
EI-REHN-1200D20171.8209−1.3079−0.11073.2395
EI-REHN-1000D20171.8370−1.3244−0.11073.2721
GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (PTB)20171.6660−1.4178−0.11073.1945
AWD-LSTM+WT+Cache+IOG (PTB)20171.7243−1.2327−0.11073.0677
Fraternal dropout + AWD-LSTM 3-layer (PTB)20171.7543−1.3914−0.11073.2564
AWD-LSTM-MoS + dynamic evaluation (PTB, 2017)20171.6784−1.4120−0.11073.2012
RNNLM + Dynamic KL Regularization20181.8910−1.2798−0.06983.2406
AWD-LSTM-MoS+Noisin+dynamic evaluation 20181.6776−1.3790−0.06983.1264
Dropout-LSTM+Noise(Bernoulli) (PTB)20181.8202−1.3848−0.06983.2747
aLSTM(depth-2)+RecurrentPolicy (PTB)20181.7427−1.3534−0.06983.1659
AWD-LSTM-MoS+PDR + dynamic evaluation (PTB)20181.6749−1.4170−0.06983.1616
AWD-LSTM-DOC (fin) (23M)20181.7192−1.3996−0.06983.1886
TrellisNet-MoS (1.4x larger)20181.7339−1.5236−0.06983.3273
Fine-tuned-AWD-LSTM-DOC(fin)20181.7170−1.2625−0.06983.0493
Multi-cell LSTM20181.8872−1.2641−0.06983.2211
Transformer-XL-ptb20191.7366−1.5731−0.13613.4458
GPT-2 (1542M)20191.5534−1.7500−0.13613.4395
AWD-LSTM-DRILL + dynamic evaluation† (PTB)20191.6937−1.4153−0.13613.2452
Adversarial + AWD-LSTM-MoS + partial shuffled20191.6629−1.3831−0.13613.1821
Tensorized Transformer (small)20191.7627−1.2641−0.13613.1629
Tensorized Transformer (large PTB)20191.7218−1.2889−0.13613.1469
R-Transformer20191.9262−1.3154−0.13613.3777
DEQ-TrellisNet20191.7566−1.4798−0.13613.3725
AWD-LSTM + DeFINE20191.7340−1.2683−0.13613.1384
bRSM + cache20192.0149−1.1840−0.13613.3351
LSTM-3-layer+Gadam20201.7692−1.3575−0.15583.2824
Tensor-Transformer(1core)+PN (PTB)20201.6776−1.2641−0.15583.0975
DiffStk-MRNN20202.0607−1.1939−0.15583.4104
ONLSTM-SYD20201.7459−1.4162−0.15583.3178
rTop-k(distributed setting)20201.9164−1.3352−0.15583.4074
3-Layer-Tensor-Transformer+AdaHessian20201.7118−1.2641−0.15583.1317
AWD-FWM (PTB)20201.7362−1.4153−0.15583.3074
CT-MoS (PTB)20201.7379−1.4153−0.15583.3090
CT-MoS + DynamicEval (PTB)20201.6760−1.4153−0.15583.2471
Selfish-RNN (ON-LSTM)20211.7468−1.3881−0.19513.3300
Selfish-RNN (SNT-ASGD) Stacked LSTMs20211.8538−1.3344−0.19513.3833
Selfish-RNN (SNT-ASGD)RHNs20211.8064−1.3492−0.19513.3507
GPT-Neo-2.7B (finetuned on PTB)20211.1673−1.8020−0.19513.1644
OPT-125M (finetuned on PTB)20221.2175−1.6814−0.23003.1288
OPT-2.7B (finetuned on PTB)20221.0334−1.7921−0.23003.0555
OPT-1.3B (finetuned on PTB)20221.0799−1.7657−0.23003.0755
Mogrifier RLSTM (PTB)20221.6325−1.3823−0.23003.2447
LLaMA-33B (LoRA finetuned)20230.8854−1.9400−0.35363.1790
LLaMA-13B (LoRA finetuned)20230.9365−1.8945−0.35363.1847
LLaMA-7B (LoRA finetuned)20230.9863−1.8722−0.35363.2122
added up over 58 rows187.661
÷ 58 rowsaPTB = 3.2355
WT103 · 86 rows
QRNN20181.5185−1.4509−0.06983.0392
4 layer QRNN (h=2500)20181.5185−1.4360−0.06983.0243
LSTM (Hebbian, Cache, MbPA)20181.4654−1.6012−0.06983.1364
Transformer (Adaptive Input Embeddings)20181.2718−1.5583−0.06982.8999
TrellisNet20181.4652−1.5236−0.06983.0586
TrellisNet-MoS (1.4x larger)20181.4652−1.5236−0.06983.0586
Transformer-XL Large20191.2625−1.5731−0.13612.9717
GPT-2 (1542M)20191.2425−1.7500−0.13613.1286
GPT-2 (762M)20191.3434−1.7252−0.13613.2047
GPT-2 (345M)20191.4211−1.6971−0.13613.2543
Transformer-XL Large + Phrase Induction20191.2405−1.4211−0.13612.7978
AdvSoft + 4 layer QRNN + dynamic evaluation20191.4472−1.4509−0.13613.0341
4 layer QRNN + dynamic evaluation20191.4997−1.4509−0.13613.0867
Tensorized Transformer (257M)20191.3263−1.5434−0.13613.0059
Tensorized Transformer (core-2)20191.2765−1.5037−0.13612.9163
Tensorized Transformer (151M)20191.2742−1.5244−0.13612.9347
All-attention network + adaptive span20191.3139−1.6244−0.13613.0744
DEQ-Transformer (Medium, Adaptive Embedding)20191.3655−1.4798−0.13612.9814
Megatron-LM (8.3B)20191.0338−1.8144−0.13612.9844
Megatron-LM (355M)20191.2858−1.7053−0.13613.1272
Sandwich Transformer20191.2514−1.6690−0.13613.0565
Compressive Transformers for Long-Range Sequence Modelling20191.2330−1.6690−0.13613.0381
Transformer-XL DeFINE (107M)20191.4103−1.5467−0.13613.0931
Adaptive LSTM + DeFINE20191.5556−1.5525−0.13613.2442
Transformer-XL DeFINE (141M)20191.3833−1.5525−0.13613.0719
TaLK Convolution20201.3674−1.6062−0.15583.1293
Turing-NLG20201.0090−1.8342−0.15582.9991
Feedback Transformer20201.2625−1.5963−0.15583.0145
TransformerXL + spectrum control20201.3655−1.4591−0.15582.9804
Tensor-Transformer(1core)+PN (WT103)20201.2529−1.5037−0.15582.9124
Segatron XL base, M=38420201.3522−1.5070−0.15583.0150
Segatron XL large, M=38420201.2330−1.6045−0.15582.9933
GPT3-6.7B (rerun of original)20200.9605−1.8243−0.15582.9406
6-Layer-Tensor-Transformer+AdaHessian20201.2989−1.5037−0.15582.9584
DeLight20201.3827−1.6012−0.15583.1398
Transformer+Recurrent Windows of Context20201.4270−1.6582−0.15583.2410
Shortformer20201.2589−1.5269−0.15582.9415
ERNIE-Doc (151M)20201.3222−1.5905−0.15583.0685
ERNIE-Doc (247M)20201.2253−1.6078−0.15582.9890
Subformer (83M)20211.3197−1.5335−0.19513.0483
Subformer (122M)20211.2989−1.5467−0.19513.0407
Subformer (96M)20211.3094−1.5384−0.19513.0430
Linear Transformer (large)20211.4983−1.5360−0.19513.2294
Linear Transformer (small)20211.5502−1.5260−0.19513.2714
SRU++ Large20211.2330−1.5731−0.19513.0013
SRU++ Large only 2 attention layers (k=5)20211.2380−1.5616−0.19512.9947
SRU++ Base20211.2625−1.5500−0.19513.0076
RFA-GATE-Gaussian-Stateful Big20211.3711−1.5574−0.19513.1236
GLM-10B-bidirectional20211.0542−1.8656−0.19513.1150
GLM-10B-unidirectional20211.0871−1.8656−0.19513.1478
Transformer-C20211.3997−1.5087−0.19513.1035
Delta RNN (+ full context)20211.5159−1.4905−0.19513.2015
DEQ-Transformer (Post-LN) + Jacobian Regularisation20211.3962−1.6078−0.19513.1992
Adaptive Input Transformer + RD20211.2570−1.6450−0.19513.0971
GPT-2 (1.5B, Curriculum Learning 45K)20211.1374−1.7169−0.19513.0494
ALiBi (L=3072, Lvalid = 3072)20211.2625−1.6739−0.19513.1315
$\infty$-former (SM)20211.2204−1.6591−0.19513.0746
PermuteFormer20211.5117−1.5277−0.19513.2346
base LM+GNN+kNN20211.2253−1.5583−0.19512.9787
S420211.3212−1.6434−0.19513.1597
GPT2+CoreLM+Fine-Tuning20211.4700−1.3633−0.19513.0284
GPT-2-Medium+Pixelfly20211.3222−1.5781−0.19513.0954
GPT-2-Small+Pixelfly20211.3522−1.5384−0.19513.0857
Gopher (7.1B)20211.0338−1.9664−0.19513.1954
Gopher (280B)20210.9096−1.8268−0.19512.9315
HSO20211.3075−1.6971−0.19513.1997
GPT3-6.7B + muP20220.9325−1.8268−0.23002.9892
Segatron-XL large, M=384 + HCP20221.2304−1.6045−0.23003.0650
Transformer Large + HCP20221.4031−1.5517−0.23003.1848
Segatron -XL base, M=150 + HCP20221.3444−1.5070−0.23003.0814
MemSizer20221.3181−1.5583−0.23003.1063
Chinchilla20220.8549−1.9631−0.23003.0480
NoPos20221.3216−1.6698−0.23003.2214
Monarch-GPT-2-Medium20221.3075−1.7053−0.23003.2428
Monarch-GPT-2-Small20221.3160−1.6756−0.23003.2215
LaMemo20221.3760−1.5591−0.23003.1651
B2T connection (16L)20221.2833−1.6070−0.23003.1203
DITTO20221.3861−1.5731−0.23003.1893
NMST+GPT-220221.3158−1.6591−0.23003.2048
Decaying Fast Weights Transformer20221.3118−1.5789−0.23003.1207
Transformer + GFM20221.3021−1.5624−0.23003.0945
Hybrid H3-355M20221.2279−1.6566−0.23003.1145
Hybrid H3-125M20221.3747−1.6186−0.23003.2233
Hybrid H3-2.7B20221.0253−1.7293−0.23002.9846
Hybrid H3-1.3B20221.0969−1.7029−0.23003.0298
Sparse Wide GPT-3 Small20231.3096−1.6483−0.35363.3116
added up over 86 rows265.053
÷ 86 rowsaWT103 = 3.0820
WT2 · 64 rows
Zoneout + Variational LSTM (WT2)20162.0039−1.3410−0.02523.3701
Pointer Sentinel-LSTM20161.9074−1.3385−0.02523.2711
AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (WT2)20171.7160−1.4451−0.11073.2718
GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (WT2)20171.6070−1.4608−0.11073.1785
AWD-LSTM+WT+Cache+IOG (WT2)20171.7135−1.2823−0.11073.1065
Fraternal dropout + AWD-LSTM 3-layer (WT2)20171.8069−1.4327−0.11073.3503
AWD-LSTM-MoS + dynamic evaluation (WT2, 2017)20171.6094−1.4575−0.11073.1776
RNNLM + Dynamic KL Regularization (WT2)20181.9385−1.3501−0.06983.3584
LSTM+Noise(Beta)20181.9186−1.4129−0.06983.4012
Dropout-LSTM+Noise(Laplace)20181.9143−1.3641−0.06983.3482
Dropout-LSTM+Noise(Bernoulli) (WT2)20181.8854−1.4129−0.06983.3680
aLSTM(depth-2)+RecurrentPolicy (WT2)20181.8096−1.3947−0.06983.2740
AWD-LSTM-DOC (fin) (37M)20181.7637−1.4162−0.06983.2496
LSTM+NeuralCache20181.8209−1.2402−0.06983.1308
GPT-2 (1542M)20191.2634−1.7500−0.13613.1495
AWD-LSTM-DRILL + dynamic evaluation† (WT2)20191.6232−1.4566−0.13613.2160
AWD-LSTM + MoS + Partial Shuffled20191.5806−1.4476−0.13613.1643
LSTM(medium)+Sememe+cell20191.9502−1.2972−0.13613.3835
AWD-FWM (WT2)20201.7899−1.4765−0.15583.4222
CT-MoS + DynamicEval (WT2)20201.6124−1.4666−0.15583.2347
CT-MoS (WT2)20201.7939−1.4666−0.15583.4162
Selfish-RNN (AWD-LSTM-MoS)20211.7997−1.4285−0.19513.4234
GPT-Neo-2.7B (finetuned)20211.0326−1.8020−0.19513.0297
GPT-Neo-125M(finetuned)20211.3416−1.6921−0.19513.2289
GPT-Neo-125M20211.5091−1.6921−0.19513.3963
GPT-Neo-2.7B20211.0565−1.8020−0.19513.0536
GPT-Neo-1.3B (finetuned)20211.0824−1.8020−0.19513.0795
GPT-J-6B20211.0366−1.8309−0.19513.0627
GPT-2 (fine-tuned with HYDRA)20211.1810−1.3451−0.19512.7212
GPT2+CoreLM+Fine-Tuning20211.5024−1.3633−0.19513.0608
GPT-NeoX-20B20220.9638−1.8797−0.23003.0734
OPT-66B20220.9703−1.9069−0.23003.1073
OPT-2.7B (finetuned on WT2)20221.0116−1.7921−0.23003.0336
OPT-6.7B20221.0358−1.8243−0.23003.0901
OPT-1.3B (finetuned)20221.0871−1.7657−0.23003.0827
OPT-2.7B20221.0959−1.7921−0.23003.1179
OPT-175B20220.9217−1.9516−0.23003.1032
OPT-13B20221.0056−1.8483−0.23003.0839
OPT-125M (finetuned)20221.2978−1.6814−0.23003.2091
OPT-1.3B20221.2151−1.7657−0.23003.2107
OPT-30B20221.0282−1.8780−0.23003.1362
OPT-350M20221.4052−1.6814−0.23003.3165
BLOOM-1.7B20221.3047−1.7813−0.23003.3160
BLOOM-1B20221.3747−1.7640−0.23003.3687
BLOOM-560M20221.4778−1.7409−0.23003.4487
BLOOM-3B20221.2448−1.8012−0.23003.2759
BLOOM-7.1B20221.1679−1.8317−0.23003.2296
Mogrifier RLSTM (WT2)20221.5798−1.4079−0.23003.2177
GPT-2+Active-SGD20231.3137−1.4451−0.35363.1124
LLaMA-33B20230.8388−1.9400−0.35363.1325
LLaMA-13B20231.1458−1.8945−0.35363.3940
LLaMA-7B20230.9773−1.8722−0.35363.2031
LLaMA-65B20230.6955−1.9648−0.35363.0139
Pythia-12b20231.0228−1.8450−0.35363.2215
Pythia-6.9b20231.0573−1.8251−0.35363.2361
Pythia-160m20231.5241−1.6905−0.35363.5682
Pythia-1b20231.2162−1.7566−0.35363.3264
Pythia-1.4b20231.1679−1.7681−0.35363.2897
Pythia-410m20231.3034−1.7243−0.35363.3814
Pythia-2.8b20231.1035−1.7929−0.35363.2500
MPT-7B20230.9983−1.8689−0.35363.2208
LLaMA-13B (LoRA finetuned)20230.7435−1.8945−0.35362.9917
LLaMA-65B (LoRA finetuned)20230.6304−1.9648−0.35362.9488
LLaMA-7B (LoRA finetuned)20230.7917−1.8722−0.35363.0176
added up over 64 rows205.628
÷ 64 rowsaWT2 = 3.2129

The year gains are the same division the other way round. Take the compute term off a row and that test's baseline off too, and what is left is that year's gain, so average one year's rows instead of one test's:

yearrows added up, log10 P − b * log10 C − atest÷ rows= gain
201540.00004g2015 = 0.0000
20168−0.20158g2016 = −0.0252
201714−1.550114g2017 = −0.1107
201822−1.535322g2018 = −0.0698
201933−4.492333g2019 = −0.1361
202026−4.050826g2020 = −0.1558
202140−7.804640g2021 = −0.1951
202241−9.429341g2022 = −0.2300
202320−7.073020g2023 = −0.3536

Both divisions land back on the numbers the fit started from, which is what the two tables were for. g2015 comes out at exactly zero rather than being set there. Every 2015 row is on the same test, so that test's baseline absorbs their level and there is nothing left for the year to carry, which is what makes every later gain a reading against 2015. Those nine are the next step.

Step 3 · each year's gain, in compute

Divide every gyear by b

The data

The 9 gyear values step 2 ended on, and the one b from step 1.

Both are in score. The answer wants compute.

The equation

compute saved = gyear / b−0.2300 / −0.0826 = 2.784

A minus over a minus gives a plus, which is right: the score fell, so the compute needed to reach a fixed score fell with it.

The result

607×

less compute in 2022 than in 2015 for the same score: 102.784.

Every year below, same division.

yearrows behind itgyear÷ b= compute saved
201540.0000−0.08260.000
20168−0.0252−0.08260.305
201714−0.1107−0.08261.340
201822−0.0698−0.08260.845
201933−0.1361−0.08261.648
202026−0.1558−0.08261.886
202140−0.1951−0.08262.362
202241−0.2300−0.08262.784
202320−0.3536−0.08264.280

Step 4 · the line through those savings

How steeply the saving rises is the answer

The data

x is the years since 2015, y is that year's saving from step 3.

The line starts at zero in 2015, because nothing had been saved yet. Only the steepness is unknown.

The equation

βE = Σ x*y / Σ x²βE = 89.43 / 204

Σ means "add up the column". Both columns are worked out below.

The result

0.438

powers of ten of compute saved, every year

yearx = years since 2015y = compute savedx * yx²
201500.0000.0000
201610.3050.3051
201721.3402.6804
201830.8452.5349
201941.6486.59016
202051.8869.42825
202162.36214.16936
202272.78419.48549
202384.28034.24264
added up89.43204
0.02.44.72015 · built from 4 results · saved 0.000 powers of ten20152016 · built from 8 results · saved 0.305 powers of ten20162017 · built from 14 results · saved 1.340 powers of ten20172018 · built from 22 results · saved 0.845 powers of ten20182019 · built from 33 results · saved 1.648 powers of ten20192020 · built from 26 results · saved 1.886 powers of ten20202021 · built from 40 results · saved 2.362 powers of ten20212022 · built from 41 results · saved 2.784 powers of ten20222023 · built from 20 results · saved 4.280 powers of ten2023
Each dot is one year's saving in powers of ten, drawn bigger where more results stand behind it. The line starts at zero in the first year, and how steeply it rises is the answer.

Not very, and the range is worth knowing

What the answer would be if slightly different papers had been written

0.438 is one number worked out from one table, and nobody designed that table. It is the results that happened to get published and collected. A handful more in one year, or a few fewer in another, and the same arithmetic comes back with a different number. How different is worth knowing before anyone quotes this one.

There is no second table to check against and no way to get one, so the 208 rows have to stand in for the wider pool they came from. Draw 208 of them at random, putting each one back before drawing the next, and what comes out is a table the same size with a slightly different mix: some rows twice, some missing altogether. That is the nearest thing available to another 208 results that might have been published. Run the whole calculation on it and it gives its own answer. That is one draw:

rows drawn

208

same size as the real table

different rows in it

126

82 never got picked

its b

−0.0778

the real one is −0.0826

its answer

0.549

the real one is 0.438

That second card is why each row goes back in before the next is drawn: without that, a draw would hand back the same 208 rows every time in a different order, and every draw would give the same answer. The draw is the same size as the real table for the matching reason, that a smaller one would wobble more for a reason that has nothing to do with the data.

Do that 400 times and count where the answers fall. Each bar is one band of answers, and its length is how many of the 400 landed in it:

8 draws fell below this range and 8 above it, the furthest at −1.73 and 1.03. A draw that happens to leave one year nearly empty can return a wild rate. They are counted here rather than drawn, because bars wide enough to reach them would leave the shape above as a single column.

Sort those 400 answers smallest to largest and the range is two positions in the list. The 10% mark sits one tenth of the way along, the 90% mark nine tenths, and each is read between the two answers either side of it:

position in 400
40.9
the one below
#40 · 0.2411
the one above
#41 · 0.2416
halving time
15.0 months
position in 400
200.5
the one below
#200 · 0.4299
the one above
#201 · 0.4299
halving time
8.4 months
position in 400
360.1
the one below
#360 · 0.6637
the one above
#361 · 0.6644
halving time
5.4 months

So eight draws out of ten land between 0.242 and 0.664, a halving time anywhere from 5 to 15 months. The middle draw comes to 0.430, close to the 0.438 the real data gives. The average of all 400 is lower, at 0.409, because the few wild draws pull it down; the middle one ignores them, which is why the range above is quoted from positions in the list rather than from an average and a spread.

So treat 0.438 as the middle of a fairly wide range, not as a precise measurement. Epoch AI, who published the data and fit it a more careful way, report a halving every 8.4 months with a range of 4.5 to 14.3. Our answer and theirs sit inside each other's ranges, which is about as much agreement as this kind of estimate allows.

That range covers one thing only: how much the answer moves when the rows move. It does not ask whether the equation is the right shape, whether three language tests speak for robotics or driving, or what the years too thin to use would have done to it. Those are larger doubts than this one, and they are the four notes below.

Four things worth knowing before quoting this

The people who published the data do it a harder way
Epoch AI split the compute into two parts, the size of the model and the amount of text it read, because those are not interchangeable. They looked at the simpler version used here and set it aside as too crude. Our answer landing near theirs is agreement, not proof.
It does not change any date on this site
This rate is added to the rate at which training runs grow, and the capability map is then fitted against that total. Make this number bigger and the map's slope flattens by exactly as much. The forecast comes out the same either way.
The early years are missing, and they matter
2012 to 2014 hold too few results to use. This is every row they have, and the gain each year would be credited with: 20121 row0.00020132 rows1.06420141 row1.461 4 rows in total, and they would carry 3 of the 12 points the line is fitted through. Counting them gives 0.549 instead of 0.438, which is a halving every 6.6 months instead of 8.2. A single 2014 model would be setting that year's whole gain. Filling those years is the most useful thing anyone could add.
It is about language models, applied everywhere
The results behind it are all language tests. The forecast uses the same rate for every domain, because nobody publishes a separate one for robotics or driving. That is an assumption, and it is written down here rather than buried.