The two-mile klystron gallery of SLAC’s linear accelerator in California, where Lance Dixon works, pictured in 2011. Image: Steve Jurvetson / Wikimedia Commons, CC BY 2.0, cropped

Claude has pushed a famously hard physics calculation further than any human team has managed. Working largely on its own for days, it computed a “nine-loop” scattering amplitude, one loop past the previous record, and a leading physicist has checked the answer, Anthropic announced on its science blog on Friday:

What’s a “loop”?

Physicists predict how particles collide and scatter using formulas called scattering amplitudes. They build the answer up in layers of ever finer corrections called loops. Each loop makes the prediction more precise, but the work grows exponentially, which is why most calculations stop at two or three.

The record was eight loops, set by Lance Dixon of SLAC National Accelerator Laboratory and Stanford and his collaborators. It was done in planar N=4 super-Yang-Mills, a simplified “toy” version of particle physics that theorists use to develop new techniques. It isn’t the real world, but methods proven there often carry over.

How Claude did it

Last month, physicist and science writer Matt von Hippel, who blogs as 4 gravitons, challenged AI to go past eight loops on the kind of compute budget a university researcher could afford. Anthropic’s team gave Claude, running as Fable 5.1 inside its Claude Science research platform, a single prompt describing the problem and then mostly left it alone, checking in with messages such as “I’m going to sleep and won’t be available for another several hours. Keep working on this until I tell you to stop.”

Claude solved it two ways: with the “bootstrap” method Dixon’s group developed, which works a bit like Sudoku by starting with every possible answer and ruling out those that break known rules, and with an indirect approach through a related, simpler quantity. According to the post, the whole project cost about $1,000 to $2,000, and the bootstrap part alone took about $100, the equivalent of 96 CPUs running for a week.

Dixon, who checked the result, was struck less by the size of the job than by how easily it could have gone wrong:

I was really quite impressed that Claude could do it directly. Not so much because it was a big computational task, but because the whole setup is very fragile.

Lance Dixon, SLAC and Stanford

Claude wasn’t the only AI on the case. A team at the Chinese Academy of Sciences led by Song He reached most of the same result independently, using GPT-6 to help with some of the steps, and published its data on September 17.

What it isn’t

Von Hippel is careful not to oversell it. Claude “used known methods, with a bit more compute than people had tried to use before,” he writes. It didn’t invent new physics, and this corner of theory is a small, specialised field. His takeaway is more about how easy it is to overestimate what’s out of reach:

There is more low-hanging fruit out there than you’d expect.

Matt von Hippel, physicist and science writer

Dixon agrees the real milestone is still ahead. “The more soul-searching moments will come when large language models start to come up with new physical principles and insights before humans,” he writes.

Why it matters

It’s the second science result Anthropic has shown off this week, after Claude flagged a CRISPR-like enzyme system. Neither is a new law of nature, but both show AI doing days of careful, fragile expert work unsupervised, for the price of a laptop, and that could change which problems researchers think are worth attempting.

Sources: Anthropic, Anthropic on X, Unite.AI.

Related