The Exponential Has Arrived
An LLM built a living harness to solve some of the hardest problems in science.
Jul 2026
Prof. Matthew Schwartz returns to describe a new project in collaboration with Anthropic. Following up on Vibe Physics — where he discussed how an LLM can perform like a multitalented graduate student at 20x the speed — Schwartz now describes what happened when he gave that graduate student a toolshop. The result, a "living harness" called BootLoops, has completed calculations across a dozen scientific fields, from particle physics to biology to economics, many of which had been open problems for decades.
Everyone says AI is improving exponentially, and for the models themselves that seems true: every few months they get smarter, by benchmarks and by feel. But in science, the results have not matched the hype. The honest scorecard so far is a smattering of results here and there — a conjecture nudged, a dataset fit, some competition mathematics. The only headline most scientists can name is AlphaFold, which is a lifetime ago in AI years, and which was classic machine learning, not a language model. Where is the exponential?
It's now here. This is what it looks like.
"I know kung fu"
Last December I tried using Claude as a research assistant, and wrote up the experience in a blog post this March: performing like a strong graduate student at 20x the speed. That was true, and it was also the wrong way to use it. A graduate student, however fast, works on one thing at a time with the knowledge in one head. The breakthrough came when I stopped treating Claude like a person and started leaning into the things it has that no person has: unlimited breadth, tireless coding, real mathematical competence, and the ability to read papers, appendices, ancient source code (all of it) at machine speed.
The expertise to solve a hard scientific problem usually exists. It's just distributed — one crucial trick in the appendix of a 1996 paper, another in a Mathematica package nobody has run in a decade, a third in the head of the one person in Marseille who never wrote it down quite right. No human can hold all of it. Claude can. Point it at the literature and it comes back like Neo in The Matrix: "I know kung fu." It absorbs the accumulated technique of every subcommunity whose papers and code it touches, and it can use all of them in the same afternoon.
Claude isn't doing science end-to-end, and nothing here ran unsupervised. I chose the problems, set the standards, killed the dead ends, and vetoed more bad ideas than I can count — the AI grad student still needs an advisor, and it still tries to declare victory early if you let it. The slope has changed: the same supervision that used to buy me one calculation now buys a research program.
The second ingredient is verification. Precision calculation has a property most of science doesn't: answers can be checked, mechanically, to a punishing standard. A wrong formula agrees with an independent computation for a few digits and then drifts; a correct one agrees to as many digits as you care to compute. Every result in this project must reproduce an independent calculation — one sharing no code and no method with the derivation — to at least thirty digits, usually far more, at check-points never used along the way. There is no partial credit and no arguing with the referee. That standard lets a self-improving system improve without accumulating junk.
A harness with a life of its own
To get real work out of an LLM you put it in a harness — the scaffolding that lets it run code, read results, and decide the next move. Claude Code is a harness. What we built is a living harness: one where the menu of tools is not fixed. When the tools at hand cannot do what a problem demands, the system writes a new tool, tests it, and keeps it. The next problem starts from a stronger bench.
In practice, running it is less one conversation with an AI than a small newsroom of them: specialized long-running sessions compute, write, and package, and an independent verification session has to sign off before any result ships. The full story of how the newsroom works — including the traps and their cures — is on the technical page.
Three weeks at the frontier
Precision calculation is where a system like this can be held honest, and particle physics is where precision calculation lives. So I started where I live. The theory predictions for the Large Hadron Collider come from perturbation theory — compute the simple version of a collision first, then add quantum corrections order by order. The corrections are drawn as Feynman diagrams with more and more closed loops, and each loop adds dimensions to an integral: a first correction might need a four-dimensional integral, the next an eight-dimensional one. Simply put, each diagram is a picture of one way a collision can happen, and its integral adds up the quantum weight of every version of that picture at once. And these integrals are really really hard — not just bigger as you climb, but structurally harder. These integrals are the calculation: the precision of the entire LHC prediction is exactly the precision with which they can be evaluated. Get them wrong in the fifth digit and a signal of new physics becomes indistinguishable from a rounding error.
How much harder has a precise geometric answer. The functions a diagram can produce are governed by the geometry of the surface where its integrand blows up. The simplest diagrams live on a sphere and give logarithms and their cousins, the polylogarithms — functions we understand completely. Add masses and the geometry becomes a torus, an elliptic curve (the $y^2 = x^3 + ax + b$ curves of number-theory fame), and the polylogarithm toolkit stops working. Beyond that sit K3 surfaces and Calabi–Yau manifolds, the higher-dimensional cousins string theorists made famous, where it is provably impossible to write the answers in any fixed menu of functions. The bananas below show the ladder: the same diagram, one more massive line per rung, climbing from sphere to elliptic curve to K3 to Calabi–Yau.
The community's plan for this frontier was patient: build the elliptic theory the way the polylogarithm theory was built, over a decade or two. I know because I helped write that plan. The polylogarithmic story took about 16 years to mature; the elliptic chapter was expected to take about as long.
BootLoops computed 30 of these frontier integrals end to end, in about three weeks. 14 were blind reproductions of the hardest known results — the answers hidden until the final comparison, as validation. 16 are new to science: integrals that had been longshot dreams of the LHC theory program, the kind of thing we put in the "someday, decades from now" column. The generic-mass ice-cream cone, proposed in 1966 and unsolved since. The non-planar double box for W-pair production with the exact top-quark mass — the branch the reference computation explicitly set aside as out of reach. The five-point non-planar family that W-plus-jet predictions have been missing. Each fell in a day or two of harness time. The "someday" column emptied out week by week.
Worked example: the night it invented numkin
One story stands in for dozens. The single most expensive step in these calculations is reduction: a diagram generates millions of related integrals, and an enormous linear-algebra pass (software called Kira and FireFly, the field's workhorses) collapses them onto a few dozen independent ones. Keep two variables symbolic at once — the spacetime dimension and a momentum — and the reconstruction of the coefficients can blow up: a handful of monstrous entries run for days while everything else is long done.
Mid-campaign, at 8:25 in the evening on July 1st, the harness hit exactly that wall on a three-loop light-by-light calculation. It invented a workaround, logged and timestamped: freeze the momentum to a number, solve the now-easy one-variable problem at a dozen different numbers (minutes each), then reconstruct the exact two-variable answer from the samples by rational interpolation, with held-out points proving the reconstruction. Turn one impossible symbolic solve into many trivial numeric ones plus one exact fit. It wrote the scripts, validated them against reserved checkpoints, and filed the result under a directory it named itself: numkin, for numeric kinematics. By the next morning the trick had been generalized into a shelf tool with a documentation page, and every later calculation that hit the same wall used it. I learned the name from the overnight report. I didn't write numkin and I didn't name it; I found it on the bench the next morning, tested, documented, and already earning its keep.
Then it kept going
The plan was particle physics. The toolkit had other ideas — because the same species of integral shows up everywhere perturbation theory does.
Gravitational waves. The waveform templates LIGO matches its signals against come from expanding Einstein's equations the way we expand collider processes — and at fifth order the integrals climb onto the same K3 and Calabi–Yau geometries. The hardest integral in the published fifth-order calculation fell almost immediately to the existing toolkit: its core is governed by the same numbers Apéry used in his famous irrationality proof, and the harness produced the exact value, confirmed to 148 digits, in 90 seconds.
Jets and the early universe. A benchmark jet observable at the LHC had carried one unevaluated integral since 2019; the harness proved which geometry it lives on and closed it. Cosmology's sky-pattern integrals have started meeting the geometry collider physics knows; the one flagged in 2024 as the first to need elliptic functions is now computed in closed form, and its elliptic parts cancel, leaving three dilogarithms.
And then, fields I never planned to touch. An 85-year-old problem about random walks on a crystal — called "the final problem" of its subject, its difficulties "insuperable" — turned out to be governed by a K3 surface; the formula the field had hunted for decades provably does not exist, and the true answer lives one rung up the geometry ladder. A string-theory constant the latest study said could not be extracted was extracted, and one of its nine numbers is a genuinely new kind of number for that field. Then biology: the statistical evidence for a four-species family tree — estimated by Monte Carlo simulation for thirty years — turns out to be an exact ratio of whole numbers, computable in milliseconds, which let us grade the field's estimators against exact truth for the first time. Once that door was open, four more fields fell in a single day each: single-cell gene expression, neutral biodiversity, natural selection's site-frequency spectra, Bayesian mixture models. Same bench, new sciences.
And then economics. The workhorse model of empirical antitrust — the demand estimates that merger decisions ride on in court — has carried a famous embarrassment for a decade: on the standard benchmark dataset, careful teams kept finding different "optimal" fits, and nobody could say which were real, because nobody could evaluate the objective with an error bar. With certified arithmetic the question closes: of 22 published candidate optima, not one is a true optimum of the exact objective. Sixteen are numerical ghosts — mirages of simulation noise — six are artifacts of optimizers giving up on flat ground, and the actual minimum is a single point, now pinned to eighteen digits. The field's modern defaults check out clean; the legacy shortcuts do not, and there is now a certified ledger saying exactly which published numbers move and by how much.
Nine fields, one bench
One result per field, a sentence or two each; every claim links to its page.
- Particle colliders — dozens of Feynman-integral families at the multiloop frontier solved with certificates, from the equal-mass kite to the complete light-by-light set.
- Gravitational waves — all four exotic Calabi–Yau geometries in black-hole scattering through fifth order in Newton's constant, solved: the integral layer beneath next-generation waveform models.
- Cosmology, correlators — the one-loop, three-site wavefunction coefficient: flagged in 2024 as cosmology's first elliptic integral; computed in closed form for any site energies, its elliptic periods cancel and dilogarithms remain.
- Cosmology, large-scale structure — the loop-integral family that blocked an exact two-loop matter power spectrum since 1996, solved; the certified spectrum matches the published Monte Carlo tables at all 140 tabulated points.
- Random walks — Watson's last lattice integral, "insuperable" for 85 years: the hunted formula proved impossible, and the true value proved to be a pairing of genus-2 periods.
- String theory — the weight-eight banana modular graph function in closed form, verifying Zerbini's conjecture at the boundary of the D'Hoker–Green theorem.
- The tree of life — to our knowledge the first exact phylogenetic Bayes factor: the tree score biology estimates by simulation, computed as an exact fraction.
- Single cells — exact likelihoods for gene-expression bursting: the "95%" confidence intervals of the standard tool turn out to be 83.4% intervals, now repaired.
- Biodiversity — neutral theory solved in both its forms and held against three decades of Barro Colorado censuses: too few rare species by a known amount, composition four times faster than drift, and a minimal successor that forecasts each species' next decade.
- Population genetics — certified likelihoods for the distribution of fitness effects, with the standard software audited against an exact reference and patches provided.
- Bayesian statistics — exact evidence integrals for mixture models — and the discovery that, under uniform priors, the standard model-selection test can need about 3.5 million observations where the textbook rule says 20.
- Jet physics — the elliptic functions of energy correlators: the stuck NNLO two-point sector solved, the four-point correlator evaluated at generic angles — where the genus-2 hyperelliptic curve identified by Ma et al. is evaluated and shown to enter real QCD — and compared with CMS Open Data jets and archival DELPHI hadronic Z events (paper).
- Economics — an open-source LLM workflow that reproduces, improves, and extends published economics articles from their replication packages: 4,452 packages from five journals, discrepancies flagged in 3,460 articles or appendices, speedups of more than tenfold in 496, extensions in 923 (NBER Working Paper 35782).
Or as a wall — hover any tile for what happened; click through for the full story, including every derivation and check.
Particle collidersParticle collidersThirty frontier collision integrals solved end-to-end: fifteen blind reproductions of the hardest known results, fifteen new to science.
Gravitational wavesGravitational wavesBlack-hole spirals for LIGO: all four exotic geometries entering the two-body problem through fifth order, closed — the hardest checked to 148 digits in ninety seconds.
CosmologyCosmologyA pattern-of-the-sky integral flagged in 2024 as cosmology's first to need elliptic functions, computed in closed form for any site energies: the elliptic parts cancel and dilogarithms remain.
Random walksRandom walksWill a walker hopping at different rates along three axes ever come home? The 1939 problem declared “insuperable”: the sought formula proven impossible — and the true answer found one level up.
String theoryString theoryNine constants of a graviton-scattering building block the latest study said could not be extracted — all identified exactly; one is a genuinely new kind of number for the field.
Tree of lifeTree of lifeThe Bayesian evidence for a four-species family tree — estimated by simulation for thirty years — is an exact ratio of whole numbers. Computed, certified, and used to grade the field’s estimators against truth for the first time.
Single-cell biologySingle-cell biologyGenes fire in bursts. The burst-model likelihood behind a standard genome-wide tool, made exact and certified — turning a long-debated question about how genes switch off into a computed verdict.
EcologyEcologyNeutral biodiversity theory solved in both its forms and held against three decades of Barro Colorado censuses: too few rare species by a known amount, change four times faster than drift, and a minimal model that predicts each species' next decade.
Still growing: the same tools have since produced results in population genetics, Bayesian statistics, and jet physics; and in empirical economics, an LLM workflow now reproduces, improves, and extends published research from replication packages.
What does this mean for AI science?
Let me quantify the exponential, since that is the claim I make in the title. The perturbative bootstrap program that BootLoops rests on began, for me, with the symbol calculus in 2010. It took our community 16 years to fully develop the polylogarithmic story. The elliptic chapter was on track to take another 16. BootLoops did the core of it — plus the 16 new integrals and new applications — in three weeks of wall-clock time. This is essentially 16 years of research compressed to a month.
One conceptual shift matters more to me than any single result here. The default picture of AI science is a model that gets so smart it can just predict the science in its head: ask the question, out comes the answer, no computers, no tools, pure intellect. I think that picture is wrong. Models are computers, and computers can use computers. We know from our own history what happens when intelligence meets tools — the invention of tools was a watershed for evolution, and everything since has been intelligence amplified by instruments, not intelligence replacing them. The idea that a model should internalize its tools, folding the calculator and the algebra system and the whole workshop into its weights, gets the direction of progress backwards. The exponential in AI science is not the models alone. It's the synergy: models improving exponentially on their side, tools accumulating on ours, each making the other more powerful. That product is what BootLoops is measuring.
Back in March I wrote that Claude worked like a talented graduate student but didn't have taste. That part hasn't changed as much as you might think. Even in this throw-Claude-at-everything mode, taste and judgment were my main contribution: which problems to chase, which standards to enforce, when to stop. What has changed is where the leverage lives. The March post was a story about the model getting better, and the models will keep getting better. But a living harness doesn't depend on any one model, or on catching the scaling curve at the right moment. Swap the engine and the bench remains: the tools, the checks, the accumulated craft. The harness keeps what the collaboration learns.
A living harness is also not a replacement for human scientists. It is instead a way to flatten the playing field. The BootLoops harness is open source and free to clone and use. Any scientist — amateur or professional — can now explore precision physics and contribute to the harness. The technical skills, built from years of study, which before had separated the experts from the amateurs, are suddenly less critical; judgment is the commodity that remains scarce. Similarly, although I used a lot of tokens and compute for this project (a month of token-maxing with Opus 4.8 and Fable 5, on two big machines), a lot of that can be recycled. A cheap LLM subscription and a laptop are enough to reproduce many of the results here, or to compute new things. One can always scale up to harder problems, but there is still a lot to be understood from the ground up. I can't wait to see what the community comes up with in improving this harness, but also in creating living harnesses in other areas for other applications, beyond anything I can conceive of.
The technical account — method, toolkit, every integral with its derivations and checks, and the full story of how BootLoops was built and run — is on the main page. A gentler introduction to the living-harness idea is here, and the original launch post is here.