METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Anthropic's Fable 5.1 computes a nine-loop scattering amplitude

Anthropic researchers answered a physicist's public challenge with Claude Science. Professor Lance Dixon spent two weeks validating a calculation that took 96 CPUs a week.

Anthropic's Fable 5.1 computes a nine-loop scattering amplitude

Image: METAL

Summary

  • On September 25, Anthropic published a guest post reporting that Fable 5.1 computed the nine-loop six-particle scattering amplitude in planar N=4 super Yang-Mills theory.
  • The calculation ran two ways in Claude Science at a cost of about 1,000 to 2,000 dollars, and the bootstrap computation amounted to 96 CPUs running for a week.
  • Professor Lance Dixon validated the result, and Song He's group at the Chinese Academy of Sciences, aided by GPT-6, reached most of the same result at nearly the same time.

On September 25, Anthropic ran a guest post on its research site reporting that Fable 5.1 had computed a nine-loop scattering amplitude in particle physics. The target was the amplitude in which six particles interact in planar N=4 super Yang-Mills theory, a hypothetical test-bed theory, and a task researchers in the field had been eyeing for years as the next loop up. Professor Lance Dixon of Stanford University and SLAC National Accelerator Laboratory spent two weeks validating the result independently.

It began with a public challenge. On August 7, Matt von Hippel, a former theoretical physicist turned science writer, wrote on his blog that "if AI companies want to impress people like me," they should tackle his old field, and asked for N=8 supergravity at seven loops or N=4 super Yang-Mills at nine loops. He explained that these problems are hard not for lack of new ideas but because the amount of calculation grows exponentially, sometimes factorially, with each added loop. The core question was whether the kind of computers an academic lab uses could get over that wall.

A loop is a unit that measures how far into the complexity of interactions between particles a calculation follows. According to von Hippel, most amplitude formulas for real particle reactions have been computed only to two loops, a few to three, and even the most precise prediction in particle physics used five. The previous step, eight loops, was a result Dixon reached indirectly in 2023 with Andy Liu, by way of a neighboring formula called the form factor and a symmetry called antipodal duality.

The challenge was taken up by Anthropic physicists Liam Fitzpatrick and Siddharth Mishra-Sharma. The two ran Fable 5.1 inside Claude Science, a paid platform for scientists. They first asked Claude which problem it was most likely to solve, gave it a one-line instruction to compute the nine-loop hexagon amplitude, and then only kept telling it to continue, with messages such as "I'm going to sleep and won't be available for another several hours. Keep working on this until I tell you to stop. Give me updates every 4-6 hours."

Claude computed the same answer by two routes: the bootstrap method refined by Dixon's team, carried one loop higher, and the indirect path through the form factor. According to von Hippel, either route would have cost an end user about 1,000 to 2,000 dollars, mostly from running Claude for so long. The bootstrap calculation itself, written in Python with SymPy, took about 100 dollars of that, the equivalent of 96 CPUs running for a week.

The result page METAL checked, dated September 16, publishes the full set of computed files. The representations produced by the two routes did not disagree on a single one of the 107,053 coefficients compared, and a test in which the same programs rebuilt the published eight-loop result and compared 1,000 random terms also matched in every case. According to the result page, writing the symbol out word by word on just one symmetric surface already yields about 30 billion nonzero terms. The page also notes that the result in function form has been computed only once, with no second independent computation, and that the programs of the computation are not distributed.

Dixon, who handled the validation, said he heard the news on September 1. He wrote that the recipe is so fragile that a single small mistake makes it all collapse like a failed soufflé, and that he was struck that Claude had written from scratch even the detailed code that cannot be fully documented in a paper. He then said, "I would assert that Claude understands our 2019 and 2023 papers better than any human, aside from my co-authors." He added an estimate that Claude is probably over a million times bigger than the custom transformer model his team had been building.

Humans were almost at the same spot. A few days after von Hippel heard from Anthropic, Song He's group at the Chinese Academy of Sciences in Beijing reported that it had already obtained most of the result. The group used GPT-6-based AI only for some of the constraints, and people built the overall framework. According to reports, Song He, Jirong Jing and Xiang Li posted symbol data through nine loops to a public repository on September 17. Dixon wrote, "So now I've been scooped by both a machine and by humans plus a machine, within two weeks."

Anthropic disclosed that it invited von Hippel to write the post and compensated him for his time, and that Dixon received Claude usage credits. Dixon and Song He's group are to write the papers explaining the results. METAL has reported on Fable 5.1 decoding a cipher that had gone unsolved for 370 years, and this time the test was the ability to push a defined calculation through to the end without errors.

Through an AI engineer's eyes, the core of this result is not new physics but the reliability of long-running work. Von Hippel himself judged that Claude ran known methods with somewhat more computing than before. What was different, he said, was that it finished a week-long calculation in one shot, without human oversight, on nothing more than "keep going." He wrote that as recently as March, AI worked at the level of a student that needed a lot of hand-holding, and that "it can do this kind of thing reliably now." The biggest lesson he named is that among goals that look far off to experts, there is more low-hanging fruit than expected. The next question for research organizations is less which calculations to hand over and more who checks the results, and with what tools.

Comments