BootLoops: A living harness for the frontier of science
AI models come and go. The tools built around them can last — and anyone can help build them.
Jul 2026
Prof. Matthew Schwartz returns to describe a new project in collaboration with Anthropic. Following up on Vibe Physics -- where he discussed how an LLM can perform like a multitalented graduate student at 20x the speed -- Schwartz now discusses what LLMs can do for science at scale. This post describes his work on building a "living harness" for precision calculations in physics and related fields. The harness, called BootLoops, has already completed extraordinarily difficult calculations directly relevant to mainstream science.
What is a harness?
A large language model thinks and talks, but can't do anything on its own. It has a brain and a mouth but no hands. It can tell you what you want to hear, plan a project, recall how a technique works, and judge which line of attack looks promising. But it cannot ingest documents, run code, or look things up on the web. To get real work out of an LLM, you put it inside a harness: a scaffolding that lets the model call software, read back what came out, and decide the next move. Claude Code is a harness for Claude. Codex is a harness for ChatGPT. The AI Overview at the top of a Google search is a harness for Gemini. Nearly every capable AI system used today runs inside a harness. However, in the normal use of a harness, the menu of tools is fixed in advance. A living harness is one that can not only write its own menu, but engineer new ingredients.
The basic idea behind a living harness is that it is self-improving. The toolkit it has collected or built improves with each project, building on previous iterations. It may struggle to solve a problem, eventually finding a needed breakthrough in a tool it didn't have. Then the next project will have that tool ready to go, and can struggle with harder things. Importantly, all of this self-improvement is at inference time, after the model is fully trained and its neural-network weights frozen. The improvement accumulates outside the model, in the tools and scaffolding, compounding from one problem to the next. An LLM with a living harness is a futuristic software package that keeps extending itself, in different useful ways, on every problem it meets. Because the toolkit ecosystem is independent of the model, it is compatible with Claude or ChatGPT or with models that don't exist yet, and can be used with past and future versions of any LLM.
Two kinds of progress
When people say "AI is getting better," they almost always mean the models: bigger, newer, smarter versions announced every few months. That kind of progress belongs to a handful of companies. It costs billions, happens behind closed doors, and there is nothing you or I can do to contribute to it. When the next model arrives, the previous one's conversations are over; whatever it "learned" while talking to you is gone.
The harness layer is the opposite in every respect. It is ordinary software plus accumulated know-how: open, inspectable, forkable, fixable. A graduate student anywhere on Earth can add a tool to it. And it is durable — when a better model ships, a living harness doesn't reset. It gets a better driver. The same toolkit works with Claude or ChatGPT or with models that don't exist yet, and every tool added today will still be there, still tested, still documented, for whatever drives it next year.
So there are two layers of AI progress: the engines, built by a few, and the vehicles, roads, and maps — which can be built by everyone. Living harnesses are a bet that much of what AI ultimately does for science will live in that second layer, where the whole community can push.
One more thing makes the bet safe: a living harness must be kept honest, mechanically. A system that improves itself needs a punishing, non-negotiable standard of truth, or it will happily accumulate junk and flattery. The best problem spaces are therefore the ones where results are hard to achieve but easy to check — where a wrong answer announces itself. Precision calculation is the purest example there is: a wrong formula matches an independent computation for a few digits and then drifts; a correct one matches to as many digits as anyone cares to compute. There is no partial credit and no arguing with the referee.
BootLoops
BootLoops is a living harness for exact calculation. (The name: the bootstrap — pulling an answer up by consistency alone, no outside help — plus the loops of the quantum diagrams where the hardest integrals live.) Its territory is a family of integrals — averaging problems, roughly: "add up all the ways this can happen" — that sits underneath a startling amount of quantitative science. Colliding particles, spiraling black holes, the ripples inflation printed on the sky, magnets near their critical point, a random walker on a crystal, even the statistical evidence for one family tree of species over another: strip away the vocabulary of each field and the same mathematical species of integral is doing the work. These integrals are famously hard — some individual ones have been open problems for decades — but their answers can be checked to hundreds of digits, which is exactly the mechanical honesty a living harness needs.
The harness began by collecting the good public software these fields already had — decades of specialized packages, built by different groups, never meant to talk to each other — and wiring it into one bench. Then it started solving. Where a tool was missing, it wrote one; where a tool fell short, it fixed it and kept the fix. Working this way with Claude for about three weeks, the toolkit grew to some thirty instruments, several of which did not exist anywhere before. Each solved problem left the bench stronger for the next one, which is the whole idea.
Every claim BootLoops makes is gated the same way: the proposed exact answer must reproduce an independent calculation — one sharing no code and no method with the derivation — to at least thirty digits, usually far more, at check-points it never saw during the derivation. Results that pass ship with a standalone script so anyone can rerun the check themselves.
One toolkit, nine fields
The same instruments, pointed at nine different sciences. Hover over any tile for what happened there; click through for the full story.
Particle collidersParticle collidersThirty frontier collision integrals solved end-to-end: fifteen blind reproductions of the hardest known results, fifteen new to science.
Gravitational wavesGravitational wavesBlack-hole spirals for LIGO: all four exotic geometries entering the two-body problem through fifth order, closed — the hardest checked to 148 digits in ninety seconds.
CosmologyCosmologyA pattern-of-the-sky integral flagged in 2024 as cosmology's first to need elliptic functions, computed in closed form for any site energies: the elliptic parts cancel and dilogarithms remain.
Random walksRandom walksWill a walker hopping at different rates along three axes ever come home? The 1939 problem declared “insuperable”: the sought formula proven impossible — and the true answer found one level up.
String theoryString theoryNine constants of a graviton-scattering building block the latest study said could not be extracted — all identified exactly; one is a genuinely new kind of number for the field.
Tree of lifeTree of lifeThe Bayesian evidence for a four-species family tree — estimated by simulation for thirty years — is an exact ratio of whole numbers. Computed, certified, and used to grade the field’s estimators against truth for the first time.
Single-cell biologySingle-cell biologyGenes fire in bursts. The burst-model likelihood behind a standard genome-wide tool, made exact and certified — turning a long-debated question about how genes switch off into a computed verdict.
EcologyEcologyNeutral biodiversity theory solved in both its forms and held against three decades of Barro Colorado censuses: too few rare species by a known amount, change four times faster than drift, and a minimal model that predicts each species' next decade.
And the list is still growing: the same tools have since closed problems in population genetics (how strongly selection acts on new mutations), Bayesian statistics (exact model-selection evidence past the field's documented wall), and jet physics at the Large Hadron Collider.
Highlights
- An 85-year-old problem, settled with a twist. The last of Watson's 1939 random-walk integrals — called "the final problem" and its difficulties "insuperable" — is now understood: the formula the field spent decades hunting provably does not exist, and the true answer lives one level up the geometric ladder, where BootLoops found it.
- Simulation replaced by exactness. The central quantity of Bayesian tree-of-life inference, approximated by Monte Carlo sampling since the 1990s, is an exact fraction — and with exact answers in hand, the field's standard estimators were graded against truth for the first time. One passed; one carries a bias that more computing can never fix.
- Errors in the human literature, caught. Along the way the harness's checks flagged misprinted formulas in a standard survey of its random-walk problem — found not by suspicion but because exact arithmetic refuses to balance around a typo.
- New mathematics, not just new answers. Several campaigns ended with objects mathematicians care about for their own sake: a new kind of number in string theory's function family, a previously unknown invariant of the random-walk problem, and exact structures that connect these integrals to some of the deepest geometry known.
- All of it checkable by anyone. Every result ships with a standalone evaluator: plain Python, no special software, rerun the digits yourself.
What this means for AI science
From a personal perspective, I am excited about this project for three reasons. The first is the science. These results are genuinely interesting and useful. Much AI for theoretical science has either been dedicated machine-learning code for one particular problem, or LLMs tackling esoteric puzzles. This isn't that. This is removing blockers to comparing theory with experiment at the Large Hadron Collider, and producing closed-form expressions that evaluate in milliseconds for reconstructing the tree of life.
The second reason is that until now I didn't really know what AI science would look like. We hear a lot about AI replacing scientists — end-to-end science from first principles, researchers sent into early retirement. That's not what's happening, at least not yet. What it's doing, through BootLoops at least, is the kinds of things we would do, but faster, cleaner, and better. I can be a little more quantitative. It took the field about 16 years to fully develop the mathematical story that BootLoops's home territory rests on, and it would plausibly have taken another 16 to write the next chapter. This project took three weeks end-to-end. Any one of the sixteen new integrals, or the additional applications, could have been a Ph.D. thesis; some of the tools could have been theses too. That's roughly 25 Ph.D. theses and 16 years of research compressed into a month. That's what the exponential growth of AI science looks like. Now we know.
The third reason is the one this whole post has been building toward: a living harness is not a replacement for human scientists — it is a way to flatten the playing field. The BootLoops harness is open source and free to clone and use. The models will keep improving on the AI companies' schedule; the harness improves on ours. Any scientist — amateur or professional — can now explore precision calculation and contribute a tool, a fix, or a solved problem, and the mechanical verification standard means contributions are judged by whether the digits check, not by who sent them. That doesn't make science easy: judgment and taste are still the scarce commodities. But the technical skills that once separated experts from everyone else are suddenly less of a wall. A cheap LLM subscription and a laptop are enough to reproduce many of the results here, or to compute new things. I can't wait to see what the community builds — in this harness, and in living harnesses for fields I haven't thought of.
For the technical account — the method, the toolkit, every integral with its derivation and checks — see the main page. The original launch post, with the full physics narrative, is here.