A week after the nine-loop amplitude story, Anthropic published another guest post labs should read slowly: Claude-shaped science, by Prof. Matthew Schwartz.
Schwartz is blunt about using frontier models in real research. Getting scientific value often feels like hand-holding for work that is brilliant in spots and poorly matched to what a scientist needed. Physicists call that an impedance mismatch: two systems that each work, poorly coupled. Models are good at science. They are not scientists.
Stop fighting the model. Find the shape.
Schwartz’s productive move was not “make Claude act more like a PI.” He stopped fighting Claude and looked for problems suited to current strengths: breadth across domains, coding, mathematics and statistics, and parsing papers and data at machine speed. He calls those Claude-shaped problems. Narrower still: harness-shaped problems.
That tool layer is BootLoops, an open-source harness for exact calculations in quantitative science. Disclosure matters: during this work Schwartz has been a visiting researcher at Anthropic, and BootLoops is not an Anthropic project. It is owned and maintained by him. Point people to bootloops.ai (and GitHub via that site) for details. Do not read this note as a claim that ResearcherFlow built, runs, or owns any of it.
From elliptic integrals to other fields
The origin story starts near Schwartz’s home turf: scattering amplitudes, the S-matrix bootstrap, and a semi-numerical bootstrap. Elliptic integrals followed. The post reports 30 integrals BootLooped end to end: 15 reproductions of known results by the new method, and 15 that had not been computed before.
Then the same mathematical patterns showed up elsewhere: ecology, population genetics, economics, linguistics, plus other highlights the post skims (phylogenetics, Earth science, genomics, sunspots, Watson’s lattice integral, cosmology, mixture models). Useful as a map. Not a scoreboard we are claiming.
Technically correct is not scientifically interesting
The sentence that should stick for lab heads, paraphrased carefully from Schwartz: in almost every case he checked with experts, Claude was technically correct, but the result was not all that interesting until the expert helped steer.
In ecology, Claude recognized Rampal Etienne’s equation for Stephen Hubbell’s neutral biodiversity theory as BootLoops-shaped and solved it at scale. On Barro Colorado Island, the mix of tree species changes 4.5 times faster than neutral theory allows. James O’Dwyer was impressed technically and still expected many ecologists to shrug. He steered toward a better question: subtract the neutral prediction and study the remainder, so species-dependent biology becomes visible.
In population genetics, Michael Desai steered away from a first rare-mutation result toward correlations between pairs of mutations. That path analyzed 5.7 billion pairs of nearby mutations from the 1000 Genomes Project and found evidence for gene conversion. In economics, an AI data editor ported replication packages for 4,452 papers across five journals (~30,000 routines) into an NBER working paper. In linguistics, AccStack assembled word-stress data for 6,072 languages plus a bibliography of 160,000 phonology works.
Credit belongs to Schwartz, O’Dwyer, Desai, and the other human collaborators named in the Anthropic acknowledgements. Claude and BootLoops sit inside that loop. They do not collapse it to a prompt.
What this is (and is not) for ResearcherFlow
We did not run BootLoops. We did not produce AccStack, the Barro Colorado ecology model, the gene-conversion analysis, or the NBER data-editor paper. We are not Claude Science, and we are not “BootLoops for labs.”
What we keep arguing for rhymes with nine loops still needing a human yes:
- Frontier models can finish technical work that used to sit only with scarce expertise, especially when the problem is shaped for the harness.
- Technically correct output is not yet science other people should trust or care about. Taste still comes from humans who know the field.
- Failure modes in the post will sound familiar: declaring victory too early, grinding forever, losing the plot without a human standard for “done.”
- Labs still need the boring half after the model stops: log, blocker, weekly writeup, approve gate, disagreements worth keeping.
That is why we talk about a Week of Record: log, blockers, weekly writeup, and nothing durable until you approve it. Keep the disagreements. A science chat can draft. It cannot replace the human yes.
Schwartz’s convex-hull metaphor lands the same idea: human knowledge is jagged, and a harness can fill middle regions no particular human has occupied yet. Filling the hull is not deciding which region was worth the expedition. That decision is human-shaped science.
What changes for us (and what does not)
ResearcherFlow will use whatever models are good at scientific research. Today that is Claude Opus 5.5 through OpenRouter. We may add other strong research models later. Do not read this note as a ship announcement for BootLoops, Claude Science, Astra, GPT-6, or any harness named in Schwartz’s post.
If your reaction is “we need a smarter agent,” you are solving half the problem. If it is “Claude-shaped work and human taste and a record we can defend after the chat ends,” you are closer to how research runs.
Try a paper (DOI) free, or start a workspace:
https://researcherflow.com
Jon Marrs
ResearcherFlow
https://researcherflow.com
Not a claim that ResearcherFlow produced, validated, or operates BootLoops or any project named above. Facts are from Anthropic’s guest post by Matthew Schwartz (Oct 1, 2026), including the BootLoops ownership disclosure. Read the source for methods and acknowledgements.