Claude computes a nine-loop amplitude in N=4 super-Yang-Mills

Overview

Anthropic has demonstrated that Claude can solve frontier problems in theoretical particle physics by computing a nine-loop amplitude in N=4 super-Yang-Mills (SYM) theory. Using a specialized scientific harness called Claude Science, the model autonomously executed complex computational recipes to reach a result that had previously remained unsolved by the community, validating that LLMs can reliably handle high-precision, multi-step scientific workflows on a modest computational budget.

The Physics Challenge: N=4 Super-Yang-Mills

Scattering amplitudes are formulas used in particle physics to calculate the probability of subatomic particle reactions. These calculations are computationally intensive and are typically measured in "loops," which represent the complexity of particle interactions. Most formulas are only calculated to two or three loops; the most precise predictions in the field generally stop at five loops.

To test new techniques, physicists use "toy model" theories. N=4 super-Yang-Mills is one such model. While not a realistic description of the physical world due to its high degree of supersymmetry (each particle has four supersymmetric partners), it is paradoxically easier to calculate with, making it a primary benchmark for honing amplitudeology techniques.

Technical Execution and Methodology

Claude achieved the nine-loop computation using the Claude Science platform, a harness that applies structured rules and prompts to the Claude LLM to ensure robust scientific behavior.

Computational Approach

Claude solved the problem using two distinct methods:

  1. The Bootstrap Method: A technique akin to a Sudoku puzzle where the model identifies the expected form of the answer and iteratively eliminates possibilities based on known rules and predictions.
  2. The Indirect Form-Factor Approach: A method that leverages a related, simpler formula (the form-factor) to derive the amplitude.

Resource Utilization

Contrary to expectations that such a feat would require massive industrial compute, the task was completed on a budget accessible to academic researchers. The bootstrap calculation was implemented in Python using the SymPy package, utilizing 96 CPUs for one week, costing approximately $100 of the total project budget (which totaled one to two thousand dollars including LLM runtime costs).

Key Findings and Implications

Reliability in Complex Workflows

One of the most significant takeaways is Claude's ability to perform "one-shot" frontier calculations without constant human oversight. The process involved developing fragile code from scratch where a single error would cause the entire calculation to fail. Claude managed this autonomously, receiving only high-level instructions to "keep working" and provide updates every 4-6 hours.

Comparison with Human Researchers

While Claude solved the problem autonomously, the result was concurrent with work from Song He's group at the Chinese Academy of Sciences. However, the human group used GPT-6 as an assistant for specific constraints rather than employing the LLM to manage the overall framework and execution.

The "Low-Hanging Fruit" Hypothesis

The achievement suggests that many frontier scientific problems may be more accessible than experts believe, provided that better software engineering practices and autonomous AI agents are applied to known methods. As noted by physicist Matt von Hippel:

"My biggest takeaway is that there is more low-hanging fruit out there than you’d expect. Even when a goal is simple and well-defined, sometimes it’s going to look much less achievable to experts than it actually is."

Expert Validation

Professor Lance Dixon of SLAC and Stanford University validated the result. He noted that Claude's success was a triumph of execution, as the model had to organize computational horsepower and follow a complicated recipe precisely. Dixon observed that Claude appeared to understand the underlying research papers (specifically those from 2019 and 2023) better than almost any human except the original co-authors.

Sources