HomeArticle

Claude has cracked a major physics puzzle. The thousand-dollar challenge has exploded in difficulty, leaving the human record holder utterly stunned.

智东西2026-09-28 11:00
The calculation results of Claude are personally verified by the former record holder.

The calculation results of Claude are personally verified by the previous record holder.

Zhidx September 26 News, in the early hours of today, Anthropic announced that Claude has achieved a major theoretical physics breakthrough, running for several days at a cost of thousands of dollars with almost no human supervision, and completed the nine-loop calculation of N=4 super Yang-Mills theory.

Blog post written by Hippel for Anthropic

The previous record holder for this difficult problem is Lance Dixon from the SLAC National Accelerator Laboratory, who completed the solution of the eight-loop amplitude in 2023. Claude's calculation results are also verified by Dixon. In the blog, he stated that he previously believed that direct solution was almost impossible to achieve, so he felt extremely shocked by what Claude has accomplished.

The origin of this incident is quite interesting too. Last month, Matt von Hippel, a former theoretical physicist and current popular science writer, launched two major challenges to major AI companies, requiring AI to complete the seven-loop calculation of N=8 supergravity or the nine-loop calculation of N=4 super Yang-Mills theory.

Anthropic chose the second challenge and adopted the Fable 5.1 model under its Claude Science platform for paying researchers. The total cost of Claude's problem-solving process for ordinary users is about $1,000 to $2,000 (equivalent to about 6,713 to 13,426 yuan), of which the bootstrap calculation implemented based on Python and the SymPy symbolic library costs about $100 (equivalent to about 671 yuan), corresponding to 96 CPUs running continuously for a week.

The core difficulty of the N=4 super Yang-Mills nine-loop calculation lies in the fact that the scale of the multiple polylogarithm symbol space, unknown coefficients and constraint equations corresponding to the planar 6-point MHV amplitude expands almost exponentially with the number of loops. This will also lead to extremely high overhead for storing and solving huge linear systems, and it is difficult for existing traditional amplitude bootstrap schemes to directly complete the full nine-loop solution by brute force.

If the N=4 super Yang-Mills six-loop calculation is analogized to 1+1=2, the nine-loop calculation is not just an increase of 3 orders of magnitude, but solving a large system of equations with tens of thousands of variables, with no lookup table available and no symbolic errors allowed, its difficulty increases explosively. Adding 3 more loops to the calculation may amplify the workload, equation scale and verification difficulty by hundreds or even thousands of times.

It is worth mentioning that a few days after Hippel received the news from Anthropic, He Song, a researcher at the Institute of Theoretical Physics of the Chinese Academy of Sciences, also contacted them, saying that his team had obtained the main part of the result with partial AI assistance provided by GPT-6.

Within a week, Claude has achieved major breakthroughs in the fields of life sciences and theoretical physics successively. This Thursday, Anthropic announced that Claude had discovered a previously unknown enzyme system hidden in bacteriophage DNA.

Full nine-loop result file:

https://smsharma.io/cosmic-nine-loops/

01. With a few simple prompts, Claude completed the solution almost autonomously throughout the whole process

Hippel's original intention in putting forward this challenge was to figure out the boundary of current AI capabilities and predict future development.

He stated in the challenge that if AI companies want to impress or shock researchers, they should go to conquer the fields that the researchers have deeply cultivated in the past. Please prove that AI can use computing resources accessible to ordinary academic researchers to solve outstanding major problems in the field of scattering amplitudes, and prove that the previously recognized computing bottleneck does not actually constitute an obstacle, by completing either the seven-loop calculation of N=8 supergravity or the nine-loop calculation of N=4 super Yang-Mills theory.

In short, it is to prove whether AI can solve cutting-edge difficult problems in the field of theoretical particle physics under the condition of an affordable academic budget.

Hippel's blog post announcing the challenge

At the end of August, two physicists from Anthropic, Liam Fitzpatrick and Siddharth Mishra-Sharma, contacted Hippel, saying that they had started to work on the challenge, and then the results were verified by Dixon. On September 1 local time, Dixon began to verify the results.

Anthropic's research uses the Fable 5.1 model under its Claude Science platform for paying researchers.

The researchers first asked Claude to judge which of the difficult problems was the most promising to conquer, and then gave one prompt: "Task: Solve the six-particle (hexagon) nine-loop MHV scattering amplitude in planar N=4 super Yang-Mills theory."

After that, researchers only need to continuously issue instructions to let the model continue to operate, for example: "I am going to rest and will not be able to intervene in the next few hours. Continue to advance this calculation until I give a stop instruction. Output a progress update every 4-6 hours."

Finally, Claude completed the solution with two independent schemes: the traditional bootstrap method and the indirect form factor scheme.

During the calculation, Claude Science had almost no manual intervention at the in-depth scientific research level, and completed the whole task at one time only by "continuing to calculate". Hippel sighed that if he used 96 CPUs to run for a week to do this kind of calculation in those years, errors would almost certainly occur, and it would actually take another two weeks to debug.

At the end of the blog, Anthropic mentioned that they invited Matt von Hippel to write this article and paid him for the contribution, and also gave Lance Dixon, who independently completed the result verification, usage credits for the Claude platform.

02. Most real-world researches only advance to 2 to 3 loops, Claude's solution to the difficult problem can help polish computing tools

The difficulty of this problem starts from amplitudeology, a branch of theoretical particle physics which is Hippel's research direction.

When particle physicists predict new particles, they must complete corresponding calculations to test the predictions. They need to calculate formulas called scattering amplitudes, which use the momentum and energy of subatomic particles to calculate the probability of specific reactions occurring between particles.

If physicists can make more accurate predictions of particle reactions, they can test whether the results obtained by experiments such as the Large Hadron Collider match the theories. Once there is a deviation between the two, it means that there may be new theories, which may help to solve many major mysteries in physics, such as the nature of dark matter and the asymmetry of matter and antimatter in the universe.

However, the difficulty lies in that the solution of the scattering amplitude formula is extremely difficult, and physicists' calculations will be truncated after reaching a certain "number of loops". The number of loops is an indicator used to measure the complexity of interactions between particles. The more loops are included in the calculation, the closer the result is to real physics, but the amount of calculation will also rise sharply.

Therefore, in reality, the vast majority of scattering amplitudes have only completed two-loop calculations, and a few have been calculated to three loops.

On this basis, researchers in the field of amplitudes will develop brand-new experimental computing techniques and test them on a special "Toy Model" theory. Choosing a toy model for testing instead of directly dealing with the complex particles in the real world is to test the capability boundary of new methods.

The challenge proposed by Hippel corresponds to these two toy models. What the Anthropic team chose to tackle is the nine-loop six-particle amplitude of N=4 super Yang-Mills theory.

"Yang-Mills" is the professional name for a class of theories. The vast majority of physical phenomena in the real world can be described by it. The four fundamental forces in nature, namely the electromagnetic force, the strong nuclear force that maintains the stability of atomic nuclei, and the weak nuclear force that causes radioactive decay (such as the decay inside bananas), all belong to Yang-Mills theory.

"N=4 super" represents supersymmetry. Physicists conjecture that every particle has a "supersymmetric partner": it has the same charge but a different particle type; matter particles (such as electrons) correspond to force particles that transmit interactions (such as photons). The academic community once optimistically believed that such particles could explain dark matter, corresponding to the N=1 version of supersymmetry, while N=4 means that each particle has four supersymmetric partners.

Hippel said that although N=4 super Yang-Mills cannot be used to explain dark matter, nor can it describe any physical phenomena in reality, it is very suitable for polishing computing tools, because the delicate balance between particles allows the calculation to only process specific variable combinations, greatly simplifying the operation. He participated in the three-loop amplitude calculation during his doctoral period.

03. Claude breaks the human record that has been maintained for three years, the previous record holder says he is shocked but not frustrated

In 2023, Lance Dixon from the SLAC National Accelerator Laboratory completed the solution of the eight-loop amplitude.

The bootstrap method, the experimental technology for this kind of calculation, has a natural adaptability to AI.

To solve the amplitude with the bootstrap method, it is not necessary to enumerate all particle interactions. You only need to have a general understanding of the form of the answer, use a special symbolic system to record all possibilities in the computer, and then verify them one by one with the predictions of other existing calculation schemes, the physical rules that the answer must abide by, and other related problems that are easier to solve as constraints.

Hippel said this is a bit like Sudoku. At the beginning, all candidate numbers are filled in the grid, and then impossible options are continuously eliminated. The ultimate goal is that only one set of solutions meets all the constraints, while retaining sufficient check items to avoid errors.

At the beginning, Dixon envisioned that the nine-loop calculation would rely on more indirect means, and even might need other types of AI. In his view, if the nine-loop calculation could be calculated by directly applying the conventional bootstrap method, someone would have completed it long ago.

Therefore, Dixon said that he was very shocked when he learned that Claude could solve it directly.

Since 2023, he and his fellow researchers have planned to solve the nine-loop form factor first, and then get the nine-loop amplitude further, but he originally thought that it was almost impossible to directly solve the nine-loop amplitude following the 2023 idea.

But Claude made it. Dixon said that what shocked him was that this problem not only has a huge amount of computation, but the whole calculation process is also extremely fragile. As long as a tiny error occurs in the calculation process, the whole system will crash completely, and then a lot of time will be spent troubleshooting bugs. In addition, a large number of implementation details are not fully written into the published papers, so Claude must write the whole set of code from scratch.

He used the nine-loop amplitude result to deduce back to the nine-loop form factor to complete the verification. For a machine that directly completed the difficult problem he thought could not be solved directly, Dixon said he would not be frustrated.

Because first of all, his team itself is developing a custom Transformer model for predicting high-loop results, and secondly, a closer look at Claude's solution process shows that it fully reuses all the methods developed by Dixon and his collaborators over the years. He believes that except for the co-authors of the papers, Claude's understanding of their 2019 and 2023 papers exceeds that of the vast majority of human researchers.

04. Challenge proposer's evaluation: AI is reusing existing methods, no breakthrough new ideas are seen

Hippel finally sighed that he originally expected to see AI break through the computing barrier in an unexpected way, but the reality is that the work completed by AI is also achievable by humans themselves. All the methods used by Claude are existing mature methods, and only more computing power is invested than ever before.

For many years, researchers with computer backgrounds have been saying that as long as more software developers are hired in the field of amplitudeology, great progress can be made, and now this view has been confirmed. There are still a large number of "low-hanging fruits" with underestimated difficulty in this field. Even if the goal is clear and clear, the subjectively assessed difficulty by experts is often much higher than the actual implementation difficulty.

At present, the technological progress of AI in the field of scientific research is getting faster and faster. However, Hippel also said that he cannot be sure whether the conclusion of this achievement can be extended to other problems. The research community of the toy model is very small, and there are many participating teams in the field of scattering amplitudes in the real world, and everyone is competing to push forward the cutting edge of scientific research.

05. Conclusion: Behind the theoretical physics breakthrough achieved by large models, it is still limited to the existing human knowledge framework

The fact that large language models can fully implement a complex particle physics calculation process and schedule computing power to complete the full solution chain still has practical significance in itself.

Claude's breakthrough in theoretical physics research proves that large models can undertake the engineering execution work under this mature paradigm, allowing researchers to not be limited to lengthy and mechanical calculation links, and focus more on scheme design and physical judgment.

But it is worth mentioning that this incident also shows that AI has not jumped out of the existing human knowledge framework, and it can only implement large-scale operations of known methods.

Looking into the longer-term future, the moment that will really trigger in-depth thinking of researchers is the day when large language models put forward brand new physical principles and insights before humans do.

This article is from the WeChat official account "Zhidx" (ID: zhidxcom), author: Cheng Qian, editor: Li Shuiqing, published with authorization from 36Kr.