The Exponential Has Arrived

An LLM built a living harness to solve some of the hardest problems in science.

Jul 2026

A hand pulling a boot up by its bootstrap, next to a looping quantum diagram — BootLoops

Prof. Matthew Schwartz returns to describe a new project in collaboration with Anthropic. Following up on Vibe Physics — where he discussed how an LLM can perform like a multitalented graduate student at 20x the speed — Schwartz now describes what happened when he gave that graduate student a toolshop. The result, a "living harness" called BootLoops, has completed calculations across a dozen scientific fields, from particle physics to biology to economics, many of which had been open problems for decades.

Everyone says AI is improving exponentially, and for the models themselves that seems true: every few months they get smarter, by benchmarks and by feel. But in science, the results have not matched the hype. The honest scorecard so far is a smattering of results here and there — a conjecture nudged, a dataset fit, some competition mathematics. The only headline most scientists can name is AlphaFold, which is a lifetime ago in AI years, and which was classic machine learning, not a language model. Where is the exponential?

It's now here. This is what it looks like.

"I know kung fu"

Last December I tried using Claude as a research assistant, and wrote up the experience in a blog post this March: performing like a strong graduate student at 20x the speed. That was true, and it was also the wrong way to use it. A graduate student, however fast, works on one thing at a time with the knowledge in one head. The breakthrough came when I stopped treating Claude like a person and started leaning into the things it has that no person has: unlimited breadth, tireless coding, real mathematical competence, and the ability to read papers, appendices, ancient source code (all of it) at machine speed.

The expertise to solve a hard scientific problem usually exists. It's just distributed — one crucial trick in the appendix of a 1996 paper, another in a Mathematica package nobody has run in a decade, a third in the head of the one person in Marseille who never wrote it down quite right. No human can hold all of it. Claude can. Point it at the literature and it comes back like Neo in The Matrix: "I know kung fu." It absorbs the accumulated technique of every subcommunity whose papers and code it touches, and it can use all of them in the same afternoon.

Claude isn't doing science end-to-end, and nothing here ran unsupervised. I chose the problems, set the standards, killed the dead ends, and vetoed more bad ideas than I can count — the AI grad student still needs an advisor, and it still tries to declare victory early if you let it. The slope has changed: the same supervision that used to buy me one calculation now buys a research program.

The second ingredient is verification. Precision calculation has a property most of science doesn't: answers can be checked, mechanically, to a punishing standard. A wrong formula agrees with an independent computation for a few digits and then drifts; a correct one agrees to as many digits as you care to compute. Every result in this project must reproduce an independent calculation — one sharing no code and no method with the derivation — to at least thirty digits, usually far more, at check-points never used along the way. There is no partial credit and no arguing with the referee. That standard lets a self-improving system improve without accumulating junk.

A harness with a life of its own

To get real work out of an LLM you put it in a harness — the scaffolding that lets it run code, read results, and decide the next move. Claude Code is a harness. What we built is a living harness: one where the menu of tools is not fixed. When the tools at hand cannot do what a problem demands, the system writes a new tool, tests it, and keeps it. The next problem starts from a stronger bench.

In practice, running it is less one conversation with an AI than a small newsroom of them: specialized long-running sessions compute, write, and package, and an independent verification session has to sign off before any result ships. The full story of how the newsroom works — including the traps and their cures — is on the technical page.

Three panels: a robot that is only a head watches a scientist work alone; a robot with two arms helps at the bench; a robot with many arms runs the whole bench while the scientist cheers.
Alone, a language model can only talk. In a harness, it can work. In a living harness, it becomes a collaborator — adapting, building new tools under human guidance, and taking on the problems people actually want solved.

Three weeks at the frontier

Precision calculation is where a system like this can be held honest, and particle physics is where precision calculation lives. So I started where I live. The theory predictions for the Large Hadron Collider come from perturbation theory — compute the simple version of a collision first, then add quantum corrections order by order. The corrections are drawn as Feynman diagrams with more and more closed loops, and each loop adds dimensions to an integral: a first correction might need a four-dimensional integral, the next an eight-dimensional one. Simply put, each diagram is a picture of one way a collision can happen, and its integral adds up the quantum weight of every version of that picture at once. And these integrals are really really hard — not just bigger as you climb, but structurally harder. These integrals are the calculation: the precision of the entire LHC prediction is exactly the precision with which they can be evaluated. Get them wrong in the fifth digit and a signal of new physics becomes indistinguishable from a rounding error.

How much harder has a precise geometric answer. The functions a diagram can produce are governed by the geometry of the surface where its integrand blows up. The simplest diagrams live on a sphere and give logarithms and their cousins, the polylogarithms — functions we understand completely. Add masses and the geometry becomes a torus, an elliptic curve (the $y^2 = x^3 + ax + b$ curves of number-theory fame), and the polylogarithm toolkit stops working. Beyond that sit K3 surfaces and Calabi–Yau manifolds, the higher-dimensional cousins string theorists made famous, where it is provably impossible to write the answers in any fixed menu of functions. The bananas below show the ladder: the same diagram, one more massive line per rung, climbing from sphere to elliptic curve to K3 to Calabi–Yau.

The banana ladder: the one-loop bubble, the two-loop sunrise, and the three- and four-loop bananas — the same topology gaining one massive line per loop order.
The geometry ladder. Each rung up demands mathematics the previous rung never needed — and past the second rung, no complete toolkit existed.

The community's plan for this frontier was patient: build the elliptic theory the way the polylogarithm theory was built, over a decade or two. I know because I helped write that plan. The polylogarithmic story took about 16 years to mature; the elliptic chapter was expected to take about as long.

BootLoops computed 30 of these frontier integrals end to end, in about three weeks. 14 were blind reproductions of the hardest known results — the answers hidden until the final comparison, as validation. 16 are new to science: integrals that had been longshot dreams of the LHC theory program, the kind of thing we put in the "someday, decades from now" column. The generic-mass ice-cream cone, proposed in 1966 and unsolved since. The non-planar double box for W-pair production with the exact top-quark mass — the branch the reference computation explicitly set aside as out of reach. The five-point non-planar family that W-plus-jet predictions have been missing. Each fell in a day or two of harness time. The "someday" column emptied out week by week.

Worked example: the night it invented numkin

One story stands in for dozens. The single most expensive step in these calculations is reduction: a diagram generates millions of related integrals, and an enormous linear-algebra pass (software called Kira and FireFly, the field's workhorses) collapses them onto a few dozen independent ones. Keep two variables symbolic at once — the spacetime dimension and a momentum — and the reconstruction of the coefficients can blow up: a handful of monstrous entries run for days while everything else is long done.

Mid-campaign, at 8:25 in the evening on July 1st, the harness hit exactly that wall on a three-loop light-by-light calculation. It invented a workaround, logged and timestamped: freeze the momentum to a number, solve the now-easy one-variable problem at a dozen different numbers (minutes each), then reconstruct the exact two-variable answer from the samples by rational interpolation, with held-out points proving the reconstruction. Turn one impossible symbolic solve into many trivial numeric ones plus one exact fit. It wrote the scripts, validated them against reserved checkpoints, and filed the result under a directory it named itself: numkin, for numeric kinematics. By the next morning the trick had been generalized into a shelf tool with a documentation page, and every later calculation that hit the same wall used it. I learned the name from the overnight report. I didn't write numkin and I didn't name it; I found it on the bench the next morning, tested, documented, and already earning its keep.

Then it kept going

The plan was particle physics. The toolkit had other ideas — because the same species of integral shows up everywhere perturbation theory does.

Gravitational waves. The waveform templates LIGO matches its signals against come from expanding Einstein's equations the way we expand collider processes — and at fifth order the integrals climb onto the same K3 and Calabi–Yau geometries. The hardest integral in the published fifth-order calculation fell almost immediately to the existing toolkit: its core is governed by the same numbers Apéry used in his famous irrationality proof, and the harness produced the exact value, confirmed to 148 digits, in 90 seconds.

Jets and the early universe. A benchmark jet observable at the LHC had carried one unevaluated integral since 2019; the harness proved which geometry it lives on and closed it. Cosmology's sky-pattern integrals have started meeting the geometry collider physics knows; the one flagged in 2024 as the first to need elliptic functions is now computed in closed form, and its elliptic parts cancel, leaving three dilogarithms.

And then, fields I never planned to touch. An 85-year-old problem about random walks on a crystal — called "the final problem" of its subject, its difficulties "insuperable" — turned out to be governed by a K3 surface; the formula the field had hunted for decades provably does not exist, and the true answer lives one rung up the geometry ladder. A string-theory constant the latest study said could not be extracted was extracted, and one of its nine numbers is a genuinely new kind of number for that field. Then biology: the statistical evidence for a four-species family tree — estimated by Monte Carlo simulation for thirty years — turns out to be an exact ratio of whole numbers, computable in milliseconds, which let us grade the field's estimators against exact truth for the first time. Once that door was open, four more fields fell in a single day each: single-cell gene expression, neutral biodiversity, natural selection's site-frequency spectra, Bayesian mixture models. Same bench, new sciences.

And then economics. The workhorse model of empirical antitrust — the demand estimates that merger decisions ride on in court — has carried a famous embarrassment for a decade: on the standard benchmark dataset, careful teams kept finding different "optimal" fits, and nobody could say which were real, because nobody could evaluate the objective with an error bar. With certified arithmetic the question closes: of 22 published candidate optima, not one is a true optimum of the exact objective. Sixteen are numerical ghosts — mirages of simulation noise — six are artifacts of optimizers giving up on flat ground, and the actual minimum is a single point, now pinned to eighteen digits. The field's modern defaults check out clean; the legacy shortcuts do not, and there is now a certified ledger saying exactly which published numbers move and by how much.

Nine fields, one bench

One result per field, a sentence or two each; every claim links to its page.

Or as a wall — hover any tile for what happened; click through for the full story, including every derivation and check.

Gallery of Feynman diagramsParticle collidersParticle collidersThirty frontier collision integrals solved end-to-end: fifteen blind reproductions of the hardest known results, fifteen new to science. Gravitational-wave chirp waveformGravitational wavesGravitational wavesBlack-hole spirals for LIGO: all four exotic geometries entering the two-body problem through fifth order, closed — the hardest checked to 148 digits in ninety seconds. CMB-like temperature speckles on the skyCosmologyCosmologyA pattern-of-the-sky integral flagged in 2024 as cosmology's first to need elliptic functions, computed in closed form for any site energies: the elliptic parts cancel and dilogarithms remain. A random walk path on a 3D latticeRandom walksRandom walksWill a walker hopping at different rates along three axes ever come home? The 1939 problem declared “insuperable”: the sought formula proven impossible — and the true answer found one level up. A torus worldsheetString theoryString theoryNine constants of a graviton-scattering building block the latest study said could not be extracted — all identified exactly; one is a genuinely new kind of number for the field. A four-species evolutionary treeTree of lifeTree of lifeThe Bayesian evidence for a four-species family tree — estimated by simulation for thirty years — is an exact ratio of whole numbers. Computed, certified, and used to grade the field’s estimators against truth for the first time. Single-cell gene expression burstsSingle-cell biologySingle-cell biologyGenes fire in bursts. The burst-model likelihood behind a standard genome-wide tool, made exact and certified — turning a long-debated question about how genes switch off into a computed verdict. Species per abundance class in the 2005 Barro Colorado census against the exact neutral expectation: the rare classes fall six- to sixteen-fold shortEcologyEcologyNeutral biodiversity theory solved in both its forms and held against three decades of Barro Colorado censuses: too few rare species by a known amount, change four times faster than drift, and a minimal model that predicts each species' next decade.

Still growing: the same tools have since produced results in population genetics, Bayesian statistics, and jet physics; and in empirical economics, an LLM workflow now reproduces, improves, and extends published research from replication packages.

What does this mean for AI science?

Let me quantify the exponential, since that is the claim I make in the title. The perturbative bootstrap program that BootLoops rests on began, for me, with the symbol calculus in 2010. It took our community 16 years to fully develop the polylogarithmic story. The elliptic chapter was on track to take another 16. BootLoops did the core of it — plus the 16 new integrals and new applications — in three weeks of wall-clock time. This is essentially 16 years of research compressed to a month.

One conceptual shift matters more to me than any single result here. The default picture of AI science is a model that gets so smart it can just predict the science in its head: ask the question, out comes the answer, no computers, no tools, pure intellect. I think that picture is wrong. Models are computers, and computers can use computers. We know from our own history what happens when intelligence meets tools — the invention of tools was a watershed for evolution, and everything since has been intelligence amplified by instruments, not intelligence replacing them. The idea that a model should internalize its tools, folding the calculator and the algebra system and the whole workshop into its weights, gets the direction of progress backwards. The exponential in AI science is not the models alone. It's the synergy: models improving exponentially on their side, tools accumulating on ours, each making the other more powerful. That product is what BootLoops is measuring.

Back in March I wrote that Claude worked like a talented graduate student but didn't have taste. That part hasn't changed as much as you might think. Even in this throw-Claude-at-everything mode, taste and judgment were my main contribution: which problems to chase, which standards to enforce, when to stop. What has changed is where the leverage lives. The March post was a story about the model getting better, and the models will keep getting better. But a living harness doesn't depend on any one model, or on catching the scaling curve at the right moment. Swap the engine and the bench remains: the tools, the checks, the accumulated craft. The harness keeps what the collaboration learns.

A living harness is also not a replacement for human scientists. It is instead a way to flatten the playing field. The BootLoops harness is open source and free to clone and use. Any scientist — amateur or professional — can now explore precision physics and contribute to the harness. The technical skills, built from years of study, which before had separated the experts from the amateurs, are suddenly less critical; judgment is the commodity that remains scarce. Similarly, although I used a lot of tokens and compute for this project (a month of token-maxing with Opus 4.8 and Fable 5, on two big machines), a lot of that can be recycled. A cheap LLM subscription and a laptop are enough to reproduce many of the results here, or to compute new things. One can always scale up to harder problems, but there is still a lot to be understood from the ground up. I can't wait to see what the community comes up with in improving this harness, but also in creating living harnesses in other areas for other applications, beyond anything I can conceive of.

The technical account — method, toolkit, every integral with its derivations and checks, and the full story of how BootLoops was built and run — is on the main page. A gentler introduction to the living-harness idea is here, and the original launch post is here.