Anthropic published a guest post today that is worth reading slowly if you run a lab: Yes, Claude can do Nine Loops.
Matt von Hippel — former amplitudeologist, science writer, and the person who issued the challenge — describes what happened when Anthropic physicists Liam Fitzpatrick and Siddharth Mishra-Sharma pointed Claude Science (with Fable) at a problem from his old field: the six-particle hexagon amplitude in planar N=4 super Yang-Mills at nine loops.
Lance Dixon at SLAC validated the result. A few days later, Song He's group at the Chinese Academy of Sciences reported concurrent progress on the nine-loop symbol, with lighter GPT-6 help on some constraints — a human-led path, not the same one-shot harness story. Everyone, von Hippel notes, has been friendly. The humans will publish and explain. Claude's role in that chapter is done for now.
That sequence matters more than the scoreboard.
What actually happened (and what did not)
Scattering amplitudes are the formulas particle physicists use to predict how particles react. Most real-world amplitudes stop at a handful of "loops." Toy models like N=4 SYM exist so the field can push methods further than the real Standard Model allows this week. Dixon had reached eight loops. Nine was the next obvious step — and, by his own account, one he expected to reach more indirectly.
Claude used known bootstrap and form-factor methods, on a budget von Hippel frames as academic-scale rather than "millions of dollars." The surprise was less a new physical principle and more that the computational and software barrier experts expected was lower than it looked. Dixon's addendum is blunt about the emotional beat: validating Claude's result also validates years of human papers Claude had to recreate in code. The deeper soul-searching moment, he writes, arrives when models invent new principles before humans do.
None of that is ResearcherFlow's calculation. We did not run Fable. We do not operate Claude Science. We are not "Claude Science for labs." Credit belongs to von Hippel for the challenge, the Anthropic team for the run, Dixon for the independent check, and Song He, Jirong Jing, and Xiang Li for the concurrent human-plus-AI path.
The lab half that does not fit in a demo
A frontier model finishing a hard computation is news. A lab absorbing that news is operations.
Someone who understands the method still has to say yes. The field still needs a durable writeup a student can inherit. Disagreements, constraints, and failed branches still need a home that is not a chat scroll. When the next model is better next month, you still need last month's record.
That is the gap we build for.
We call the portable unit a Week of Record: experiment log, blockers, a short weekly writeup, and an approve gate before anything becomes the lab's durable record. Literature in. Next experiment out. You decide what lands. A science chat can draft. It cannot replace the human yes, and it cannot replace the inherited file.
What this changes for ResearcherFlow (and what it does not)
ResearcherFlow will use whatever models are good at scientific research. Today that is Claude Opus 5.5 through OpenRouter. We may add other strong research models later. Do not read this Journal note as a ship announcement for GPT-6, Astra, or any harness Anthropic used for the nine-loop run.
What the nine-loop story reinforces is older than any one model launch:
- Frontier AI can now finish work that used to sit only with top human experts in a narrow subfield — sometimes on a budget a carefully funded group can actually spend.
- Concurrent human-plus-AI paths still matter. Song He's group is the reminder that "the model did it alone" is not the only useful story.
- Validation and publication are not optional polish. They are the difference between a prompt trail and science other people can trust.
- Labs still need a place for the boring half: the log, the blocker, the weekly writeup, the approve gate.
If your reaction to nine loops is "we need a smarter chat," you are solving last year's problem. If your reaction is "we need a smarter chat and a record we can defend after the chat ends," you are closer to how research actually runs.
Try a paper (DOI) free, or start a workspace:
https://researcherflow.com
Subscribe on this page if you want the next note when frontier AI does something a lab should actually care about.
Jon Marrs
ResearcherFlow
https://researcherflow.com
Not a claim that ResearcherFlow produced or validated the nine-loop amplitude. Facts above are from Anthropic's published guest post and Dixon's addendum; read the source for methods, cost framing, and concurrent results.