Harvard University’s Lyman Laboratory of Physics in Cambridge, Massachusetts, in 2012. Image: Daderot / Wikimedia Commons, CC0, cropped

A Harvard physicist says he has turned Claude into a paper factory, and it now works well beyond physics. Matthew Schwartz, a professor of theoretical physics, wrote on Anthropic’s science blog on Thursday that he has produced 36 manuscripts in 18 fields with 19 co-authors in three months, picked from about 400 candidate problems, using an open-source toolkit he built with Claude called BootLoops.

It is the follow-up to his earlier experiment, in which he found that Claude worked like a strong graduate student at about 20 times the speed but needed every sentence corrected. This time, he says, he stopped trying to make Claude work like a human scientist and went looking for problems that suit what it is already good at. He calls them “Claude-shaped” problems.

What BootLoops is

BootLoops started as a set of tools for scattering amplitudes, the calculations that turn particle collisions at the Large Hadron Collider into predictions. Schwartz asked Claude Fable 5 to port methods scattered across Wolfram Language, C++, Python and Julia, plus some that existed only in papers, into one framework. He says it reproduced the results of one of his own papers in about 20 minutes, where his original code had taken weeks, and then told him there was a better algorithm he didn’t know about.

Pushed onto elliptic integrals, which he describes as “really hard, even for people like me”, the toolkit computed 30 Feynman integrals from start to finish: 15 known results reproduced by a new method and 15 that had never been computed before. It is the same team behind last week’s nine-loop physics record, which used a related approach.

BootLoops is MIT-licensed and works with any model, not only Claude. Anthropic notes that Schwartz has been a visiting researcher there during the project, but that BootLoops is his project, not Anthropic’s.

From forests to sunspots

Because the same equations turn up in very different sciences, Claude kept spotting places where the toolkit would apply outside physics. Schwartz lists results in ecology, population genetics, phylogenetics, economics, linguistics, earth science, genomics, cosmology and statistics, each done with experts in that field. Among them:

  • Ecology: it solved an equation from 2005 for neutral biodiversity theory that nobody had been able to solve at scale, and found that tree species on Panama’s Barro Colorado Island change 4.5 times faster than the theory allows.
  • Genetics: analysing 5.7 billion pairs of nearby mutations from the 1000 Genomes Project, it found evidence for gene conversion, a mechanism most analyses ignore.
  • Economics: an AI “data editor” ported the replication files of 4,452 papers from five leading journals to open-source code and checked their published numbers, written up as an NBER working paper.
  • Linguistics: a database of word stress covering 6,072 languages.

Most of these are Schwartz’s own account and, he says, are still being explored and checked. The pattern was the same each time: Claude produced a result he found exciting, the field’s expert was “unmoved but sees the potential”, and together they turned it into something worth publishing.

“Claude loves to declare victory”

The post is blunt about where Claude still falls short. It has no sense of time: after three days on one project, it told Schwartz the formula was solved after “Two years of campaign”. Its time estimates were always far too long or far too short. It preferred grinding through days of calculation to building a tool that would take minutes, and it often lost track after its context was compacted.

His sharpest warning was about claims of success:

Claude loves to declare victory. “Done, with one asterisk” is often “not done at all.”

Matthew Schwartz, Harvard University

On one project, Claude was proud of a proof that rested on “one unproven lemma”. The lemma, he writes, “was the whole proof”. His advice is to set rigid standards for success, look at every plot yourself, question every conclusion and supply the taste, because Claude “can find Claude-shaped problems, but there are thousands, perhaps millions, of them”.

What it means for scientists

Schwartz is optimistic but not starry-eyed. He argues the focus on headline maths prizes risks being a “footgun” that distracts from what is already possible, and that real progress will still come step by step from real-world data. He also admits he no longer knows how to train graduate students, and worries about credit:

I hope we can figure out a way to give humans the credit they deserve even for Claude-shaped science.

Matthew Schwartz, Harvard University

The flood of AI-assisted papers is already a problem for the people who publish them: arXiv has just limited scientists to two submissions a month, and leading mathematicians have asked AI labs to change how they release AI-generated results.

Why it matters

This is one of the most detailed first-hand accounts yet of an established scientist using AI agents at scale, and it cuts both ways: dozens of results in months, but only because a human picked the problems, checked the work and brought in experts who knew what mattered. If BootLoops spreads, the bottleneck in science may move from doing the calculations to deciding which ones are worth doing.

Sources: Matthew Schwartz, “Claude-shaped science”, Anthropic Science blog, BootLoops.

Latest Anthropic news

More Anthropic news