The Einstein Test Failed. The Math Claims Are a Different Experiment

Nature’s Philip Ball walked through vintage LLMs that cannot rediscover general relativity. Axios says labs are still bragging about open math problems. Do not mix the two scoreboards.

Desk with a printed 1911 physics paper beside a laptop showing a sparse loss curve, cool overhead light, no logos

Demis Hassabis wanted a clean exam. Train a model on everything known before a cutoff — he suggested 1911 — and see whether it invents general relativity. He called that a good test for AGI at the India AI Summit in February. Philip Ball’s Nature feature, dated September 9 with a correction on the 14th, is the write-up of what happened when people actually tried versions of that idea. Early vintage models did not jump. They showed the usual limitation: today’s systems are better at grinding through known structure than at abducting a new one from almost no data.

The same week, Axios ran the other scoreboard. OpenAI says GPT-6 Astra made “substantial progress” on old math and theoretical computer science problems. Anthropic and Google reported their own math advances. Staffers at Anthropic apparently told Claude to take a real stab at the Riemann hypothesis and to believe in itself after a weak first pass. That is a lab notebook, not a 1911 time machine.

We already did Astra’s safety and hidden-reasoning story. This is the discovery claim. Keep them apart.

What the Einstein test is asking

Owain Evans, at Truthful AI, talked in December 2024 about “vintage” or historical models: train only up to a date, then ask what they rediscover. Hassabis’s version is specific. Relativity is the target because it is famous, sparse, and not a pattern you can average out of a million textbooks printed after 1916.

Tom Zahavy, also at Google DeepMind, put the objection in a January position paper Ball cites: “LLMs can’t jump.” The missing move is abduction — inventing a cause for a singular fact — not induction from piles of examples. Einstein did not have a pretraining corpus of later textbooks. He had anomalies and a willingness to throw out Newton’s stage.

Ido Kaminer’s Technion group posted a July preprint, “Can AI follow in Einstein’s footsteps?” Their answer, as Ball reports it, is not “never.” It is “not with the current recipe.” If you want a relativity-sized leap, you probably have to change how the model is built, not prompt harder.

Sendhil Mullainathan’s group at MIT gave an orbital-mechanics foundation model synthetic solar systems that obeyed Newton. The model never recovered the actual inverse-square law. It invented a wrong force law per system. That is the unsexy result hiding under AGI rhetoric. A network can fit trajectories and still miss the one equation you cared about.

Math is not relativity, and the labs know it

Ball is careful on the other side. LLMs have been useful in mathematics in ways that look like they found structure in the training distribution. A May result from an OpenAI system that disproved an eighty-year-old Erdős conjecture was, in Mullainathan’s words, a “genuine conceptual advance.” It needed small new abstractions. It also assembled ideas that already existed in the literature. That is still a result. It is not 1915.

Axios’s Astra write-up is a vendor claim about open problems. “Resolved or substantial progress” is a phrase you should refuse to launder into “proved.” If you cannot read the paper, you do not have a theorem. You have a blog post. Anthropic’s Riemann prompting is even thinner. “Believe in yourself” is not an evaluation harness.

Inception’s Mercury 2.5 was a speed claim you could at least time. Discovery is worse. There is no tokens-per-second for “did we invent a world model.” There is a paper, a proof, or there is not.

Google’s math work and OpenAI’s math work can both be real and still fail Hassabis’s test. Rediscovering a lemma inside a corpus that already contains twentieth-century mathematics is a different sport from being locked out of 1916.

How to read this week without joining a church

If you run evals, stop using “could this be Einstein” as a slide. Ball’s piece is useful because the experiment is now empirical: people trained vintage models, and the early ones did not produce GR. That is a negative result worth logging. It does not say models cannot help a living physicist. It says next-token machines are a poor reconstruction of a one-off conceptual leap.

If you ship a math agent, log three things that Nature and Axios keep mixing:

  1. Was the training cutoff honest, or did the web leak the answer.
  2. Did a human mathematician accept the write-up, or only the company’s blog.
  3. Did the model invent a representation, or search a space we already knew how to search.

The orbital-mechanics paper is the one I would hand a skeptic. Wrong law, confident fit. That failure mode will show up in chemistry and in code repair too. We already watched coding models get scored on vendor benches. Science benches will get the same disease if “progress on an open problem” stays undefined.

Hassabis can keep the Einstein test as a north star. Just do not grade Astra on it, and do not grade vintage models on Astra’s press cycle. One is a locked library. The other is a lab with the lights on and a Slack channel full of hints.

The correction on September 14 means you should read the Nature page, not a summary of the summary. I am not going to guess what they fixed. The plot that survived is enough: the jump is still the hard part, and math progress, even when real, is not that jump.