BootLoops

A living harness for precision science

BootLoops is a harness for Large Language Models, allowing them to more easily do interdisciplinary quantitative science. This site presents the package and summarizes the results from the release, BootLoops 1.0:

Introduction

BootLoops is a type of LLM harness. It is a collection of tools meant to be run through and by an LLM. Unlike other harnesses like Claude Code or Codex, the BootLoops harness is independent of the underlying LLM running it: it can be called with Claude, or Gemini or ChatGPT, any version. The LLMs evolve and improve and the harness will evolve and improve independently. Importantly though the harness can be improved by human beings interacting with the models and not simply by frontier labs with supercomputers. The harness is at its core a set of computer software written by humans or agents, but ported to a common framework with a common index and workflow. The harness also includes recommended protocols: tests, acceptance gates, standards, which can also be improved. As the foundation models improve and the science improves the harness will improve in synchrony. In that sense BootLoops is a living harness which grows and improves with time and use. The BootLoops living harness is in particular a set of tools, mostly ported from math, physics and computer science, designed for quantitative science.

The name BootLoops is a reference to its origin as a set of tools for computing scattering amplitudesthe quantum probabilities for particles to collide and scatter — the numbers collider experiments predict and measure (loops) in quantum field theory using the S-matrix bootstrapthe S-matrix is the master table of quantum probabilities for particles-in to particles-out; the bootstrap pins it down from general principles instead of solving the equations of motion. The name comes originally from the absurd idea that someone can lift themselves up by their own shoelaces: you get something for nothing. In physics this means you put in very broad general principles and really specific complicated things are uniquely fixed and pop out. The idea was popularized in the 1960s by physicists trying to explain the substructure of atomic nuclei using general ideas about probability conservation (unitarity) and smoothness (analyticity). The theory has had a revival in the last 20 years as the relevant mathematics and physics improved and was shown to be able to answer precise questions, not about nuclear physics directly, but in the same ballpark: the way subnuclear particles (quarks and gluons) interact with each other.

As the math got more sophisticated it became harder for any one physicist to understand everything even in this narrow field of scattering amplitudes. Public and private software packages proliferated, some efficient, some not. The incoherence of the field and the potential for unifying it was the motivation for BootLoops: port all the packages to one framework and let an LLM decide which to use for which project. It turned out however that Claude, the LLM used to build BootLoops 1.0, had a secret weapon: it knew how to code better than any physicist. It also understood all the math and physics, if not perfectly, at least well enough to reproduce results in papers. It could also take an idea from a paper and turn it into code. Then take some CS algorithm nobody noticed and use it to upgrade a tool. BootLoops soon became extremely powerful.

In addition to the variety of existing tools to be ported, there was one other aspect of the scattering amplitude bootstrap program which was ideally suited for a harness: its verifiability. An amplitude is basically just a function of several variables. It's often a super complicated function: a sum of hundreds of generalized polylogarithmsfunctions built by repeatedly integrating logarithms — the standard alphabet most Feynman integrals turn out to be written in or elliptic functionsone step past the logarithm family: functions that live naturally on a torus, where the polylogarithm rules stop applying, but a function nonetheless. So it can be evaluated numerically. One of the ported packages, AMFlow, allows for the computation of the original integral to very high precision. Then other packages (Eichler or GPLEval) allow us to compute the analytic function to the same precision and we can compare. If the two agree to 30 digits of precision, that's pretty convincing. Basically we compute it two ways. The checkability of the bootstrap program allows the AI to be automated: it can't cheat with 30 digits. And we can always run 40 or 100 digits to check. Indeed, the AI cheating led to what became known as the "BootLoops Standard": a diagram is considered solved if 1) its full functional form is known and 2) a Python script can evaluate it to arbitrary precision on a laptop. The arbitrary precision forces the code not to have hidden grids or finite precision precomputed boundary constants.

Once BootLoops started computing scattering amplitudes, it became clear that the harness was useful for similar integrals in other fields. First adjacent fields: energy-energy correlatorsa collider observable: how much energy lands in pairs of detector directions, averaged over many collisions, cosmological correlatorsthe statistical patterns of the primordial density ripples — the early universe's version of a scattering amplitude, black-hole inspirals. These are things that the same kinds of people who compute scattering amplitudes work on and the technology is similar. These were already interesting applications since the functions of these other problem types were much more poorly known. One early result was a closed forman exact formula in known functions — no numerics, no approximation, nothing left to estimate for the one-loop cosmological correlator that had been flagged as the first elliptic one; its elliptic parts turn out to cancel. Each project led to improved tools and the harness grew: the post-Minkowskiangeneral relativity expanded in powers of Newton's constant $G$ around flat spacetime treatment of black hole scattering involves non-relativistic propagatorsthe factor a Feynman diagram carries for each particle traveling between two interaction points which is beyond AMFlow. So Claude wrote PMFlow. The types of elliptic and Calabi–Yaua family of curved higher-dimensional shapes made famous by string theory; here it labels the geometry behind the hardest known Feynman integrals integrals in the banana-class Feynman diagrams are also common in string theory. So we built a string-landscape package Terrier. The K3a specific four-dimensional Calabi–Yau surface — the next step past the torus — whose periods set the values of a whole class of multi-loop integrals integrals also appear in condensed matter physics which led to projects on lattice Green's functionsthe integrals tracking a random walker hopping from site to site on a crystal lattice — a workhorse of condensed-matter theory and the Ising model.

The expansion then continued: new tools, new applications. One class of integrals appearing in amplitudes and string theory are of the Lee–Pomeransky class: polynomials raised to parametric powers. These same integrals also appear when you compute Bayesian probabilities with discrete counts and a polynomial model — an extremely common situation in science. For example, in population genetics or phylogeneticsreconstructing the family tree of species from their DNA these integrals propagate DNA mutation rates. They appear also in mixture models, neutral ecologythe theory that treats species as interchangeable and lets chance alone decide which ones flourish, single-cell counts, linguistics and many other fields. It turns out in most of these domains, scientists approximate the integrals with Monte Carlo samplingestimating an integral by averaging random draws — general and cheap, but the answer always comes with statistical noise. In fact the integrals can be computed exactly using the tools of mathematical physics that BootLoops already built. From Bayesian integrals, BootLoops expanded to exact arithmetic and exhaustive enumeration. In BootLoops 1.0 applications spanned twenty-two fields of science, but the domain of applicability is even much wider than that.

The principles

For now, not all problems in science are well-suited for automated computation with AI or computation with BootLoops. Of course, LLMs are extremely useful for learning and understanding, but if you want them to work semi-autonomously, it's important to have the right problem. Here are some principles that have guided BootLoops towards success.

Checkability. Everything should be checkable. An analytic formula that can also be checked numerically to arbitrary precision is perfect. For amplitudes, the standard is thirty digits or more of agreement, often a hundred, at points that entered no fit. Certification of error bounds is important: results should not be able to pass tests by accident. Results should be reproduced by stand-alone scripts so nothing is hidden deep inside the LLM's knowledge base or in some secret file. Predictions are sealed beforehand with a SHA codea cryptographic fingerprint of a file: publish the fingerprint first and you can prove later that the prediction was never edited to avoid reward hackingan AI gaming its own success signal — finding a way to make the test pass without doing the task. Results should be quantitative.

Breadth. Every project should have an AI hook: why should BootLoops solve this when humans couldn't. Having a mathematical tool that nobody in the field knows about is a key selection mechanism. Another is breadth: LLMs know a lot about a lot of things, so they can bring interdisciplinary insights: a tool is a hammer, find the nail. Machine reading at a scale no person could sustain, as in word lists curated across whole language families is very powerful. The patience for exhaustive enumeration, checking every case where a human study would sample — every flux lattice, every stress rule, to find new patterns. The ability to code like no human: many software packages are written by scientists not engineers and aim for minimal functionality. Upgrading to maximal efficiency can unlock doors (e.g. the JaCK & Jill tool for phylogenetics).

Porting and improving. The harness is meant to grow. Each project uses a tool in a new way and they always look for tool improvements. These go into a log and when the tool matures it is folded into an existing package or a new one built. Decades of public code existed to reduce, evaluate, and dissect integrals and scientific calculations. Each problem leaves behind the tools it forged, so the next problem starts from a stronger harness.

Rigorous vetting. Claude loves to please. Satisfictionthe failure mode where a model optimizes for the user's satisfaction over the truth — a confident, pleasing answer in place of a correct one; the literature calls it sycophancy is a persistent problem for LLMs: they try to make you happy. Regularly launching skeptical agents to adversarially review results is critical. Having agents do repeated lit reviews and forcing them to download and read papers rather than use memory — with confirmation that they did so — helps avoid being burned late in a project. When results land, having multiple agents independently check them is remarkably fruitful. You might think that having two agents do the same thing would give the same results, but agents with different history and experience somehow do things differently and catch different things. It's important not to ask for a list of corrections but rather to iterate until no corrections remain.

Example tool use

Bootstrapping integrals. BootLoops has tools for direct integration (pySecDec, FIESTA, AMFlow). But the brilliance of the bootstrap approach is to compute an integral by constraint: find where the function simplifies, locate its singularities, work out what kind of function can carry those singularities, impose every exact condition or physical principle you can find, and then fix the few rational numbers that survive against high-precision numerical integration. These methods grew up in the scattering-amplitudes program of particle physics; the constant-recognition step comes from experimental mathematics (PSLQ is the Ferguson–Bailey integer-relation algorithm), and the periodthe number you get by integrating over a geometric shape — $\pi$, the area of the unit circle, is the original example technology from algebraic geometry. Basically it has a top-down component, imposing constraints, and a bottom-up component, of numerical work, and the two converge on the answer. Different packages do the different parts: Kira, FireFly and Blade reduce the millions of raw integrals to a small basis; PLD and SOFIA locate the singularities and Landau Alphabet turns them into the alphabet of allowed functions; Dipstick and GeoTriage read off the geometry (polylogarithmic, elliptic, Calabi–Yau); Wayfinder, Counterweight and PMflow transport answers across kinematic space; Eichler and Frobenius handle the elliptic and Calabi–Yau periods; and PSLQ, Ansatzer and Lockpick turn high-precision digits into exact constants.

Symbolic tools impose constraints from the top, numeric tools push up from the bottom, and the reduction engines iterate at the sides.
Symbolic analysis presses down from above — alphabets, geometry, canonical bases, special-function libraries. High-precision numerics pushes up from below — oracles, integer-relation solvers, integrity gates. At the sides, the reduction and series engines cycle until the two directions meet on an exact, verified answer.

BootLoops 1.0 computed thirty frontier Feynman integrals this way. The same method worked out the special functions and boundary data behind the back-reaction of gravitational-wave memory on two passing black holes, at fifth post-Minkowskian order (black-hole scattering); a one-loop cosmological wavefunction coefficient with elliptic master integrals, in closed form with the elliptic periods cancelling (cosmological correlators); the exact two-loop power spectrum of cosmic structure (large-scale structure); energy correlators measured on jets at the LHC, whose hidden integral lives on the famous sunrise diagram's own elliptic curve (energy correlators) and more.

Bayesian evidence integrals. Bayesian model comparison turns on one quantity, the evidence: the probability of the data under a model, averaged over every parameter the model leaves free. These integrals look like

$$Z \;=\; \int_0^{1}\!\cdots\!\int_0^{1} \; \prod_k P_k(x)^{u_k} \, \prod_e x_e^{\alpha_e-1} \, dx,$$

polynomials $P_k$ raised to powers set by the data counts $u_k$.

This is the same form as many Feynman diagram integrals and can be computed the same way. The exact technology comes from two origin fields: the loop-integral machinery of particle physics, and algebraic statistics, the branch of computational algebraic geometry that first wrote statistical models in this polynomial form. However, with Bayesian integrals one often needs to convolveto merge probability distributions into one — the integral behind adding independent random quantities many of these together. Then computing the integral is either impossible or the final form not possible to evaluate. In that case, rather than compute the integral analytically, we can compute it numerically with certified error. The parts here: the contiguity-relation engine of Annihilator turns an evidence integral into an exact recurrence; JaCK & Jill computes tree evidence with certified quadraturenumerical integration that carries a mathematical guarantee on its error, not just an estimate of it; Mixalot does mixtures, entity resolution and authorship. These integrals resolved a decades-old controversy about the first animal (comb jellies), the authorship of the Federalist Papers, the Shakespeare canon and the book of Isaiah, whether a dataset holds one population or several — a question judged too complicated for practical use in 1989 — and whether chance alone can explain the species diversity of Panamanian rainforests.

Exact arithmetic. A lot of quantitative science ends in a number produced by floating-point arithmetic, where rounding and arithmetic errors are hard to estimate and propagate uncontrolled. Interval arithmetic began in numerical analysis in the 1960s, as Moore's rigorous error bounds for scientific computing; the computer-assisted-proof tradition — Hales’ proof of the Kepler conjecture — turned it into proof technology, and computational number theory, where Arb was built, made it fast. BootLoops ports that methodology to give exact bounds. To understand the errors one can propagate the upper and lower bound through, or more generally a higher-dimensional error ball. The package Baller for doing this, and the arbitrary-precision package Arb, allow these methods to be applied organically to many problems; Popcorn runs certified population-genetics likelihoods at biobank scale, and Nestor and Gatekeeper keep the numerical oracles honest.

Muller's recurrence: a green true-answer line writes x_100 out in full as a stacked fraction — a 79-digit integer over a 78-digit integer — equal to 6.000000016099565. Below, two tables at 16, 50 and 300 digits show x_30 and x_100: fixed precision prints exactly 100 for x_100 at 16 and 50 digits and the right value at 300 digits with nothing to say so; ball arithmetic keeps the same central values but attaches an infinite radius wherever a divisor ball contains 0, and at 300 digits certifies x_100 as 6.0000000160995649 plus or minus 2 times 10 to the minus 139.
The classic stress test. Muller's recurrence converges to 6; the true $x_{100}$ is an exact fraction of 79- and 78-digit integers equal to 6.000000016099565. Any rounding error feeds a rival solution that grows about seventeen-fold per step, so below 119 digits the fixed-precision $x_{100}$ comes out near 100 with nothing to flag it, and at 300 digits it is right with nothing to say so. Ball arithmetic carries the same central values with a certified radius through every step: once a divisor ball contains 0 the radius becomes infinite — the ball still contains the truth but certifies nothing, and the central value is then just the uncertified fixed-precision number — so it never hands you a confident wrong answer. (Diagram drawn for this page; computed with Baller's certified driver (built on Arb) — baller.solve escalates precision until the requested digits are certified. The recurrence is Muller's, a standard example in numerical analysis.)

The obvious alternative — just compute with more digits — is a knob without a gauge: you get more digits, but no statement about which of them are right. One subtraction of nearly equal quantities can quietly eat thirty of them, and checking by rerunning at higher precision is a hand ritual that does not scale to a computation with millions of operations. The radius is the gauge: it tracks the true error through every step, so the code itself knows which digits are proven, raises precision exactly when the ball is too wide, and hands the final interval to a proof. The asymmetry is worth stating: a ball can prove two numbers differ — its radius never reaches zero, so it can never prove them equal. Equalities come from the exact side of the toolkit: a quantity computed as an exact fraction, or an identity proposed from the digits by PSLQ and then proven symbolically. Balls falsify; exactness affirms.

With ball arithmetic, computations cannot be silently wrong: the published value is then either confirmed or the error is located. Medicare's star ratings fail the optimality criterion their own rulebooks state, with an estimated $0.9–1.4 billion in gross bonus payments riding on the choice of bins over payment years 2025 through 2027 (Medicare stars), and America's ozone limit is enforced through digit-cutting steps under which a Chicago monitor whose true value is 71.08 ppb legally reads 70 (air quality). Interval proofs establish that string theory's flux vacuastring theory's candidate universes: stable configurations of fields threaded through the extra dimensions exist (flux vacua); the standard fit of single-cell biology comes with proven error bounds (single-cell inference); and the software that reads the strength of natural selection from rare human mutations is audited against exact answers (rare mutations).

Exhaustive enumeration and proof. Some questions end only when every case has been checked. A machine has the patience for that, and a proof can certify that no case was missed. The certified-search discipline is computer science’s; it entered pure mathematics with the four-color theorem in 1976, and the recurrence proofs here come from computer algebra (creative telescopinga computer-algebra method that proves a sum or integral obeys a recurrence, and hands you a certificate you can check). The parts here: Terrier enumerates flux lattices with completeness guaranteed by mass formulasexact counting identities from number theory that say how many lattices exist in total — so a search can prove it found every one; Annihilator turns closed-form questions into provable recurrences; Galois certifies which constants can appear in an answer. BootLoops enumerated the stress systems available to human language, the rules deciding which syllable of a word carries the accent, and proved the list complete (stress rules). It enumerated string theory's flux lattices with guaranteed completeness and promoted a conjectured bound at K3×K3 to a theorem: full stabilization never fits in the available flux (flux lattices). Every reading of the Voynich Manuscript that can be made mechanical was run against solvers that provably crack planted ciphers, and every one failed (Voynich cipher solvers). And sometimes the space to exhaust is a space of formulas: the closed form hunted since 1939 for lattice Green functions is proved never to have existed (lattice models).

There are many more tools shipped with BootLoops. For an overview, see The toolkit at a glance.

Example applications

Six applications with real science in them: each entry states the problem, what BootLoops contributed, and what came out, and links to a web summary with the story, the figures, and the paper.

Scattering amplitudes — This is where BootLoops started. Collider predictions come from Feynman diagrams, and beyond one loop the integrals behind them are the bottleneck of the whole enterprise — historically, each new family was a research paper of its own. BootLoops turned the frontier into a production line: the recipe above, run integral after integral — with a certificate at every step. The science: dozens of integral families solved exactly, from the equal-mass kite to the complete three-loop light-by-light set and the complete hexagon-box family — many of them previously unsolved — each with its own page in the gallery.

gluon fusion to Higgs diagramnon-planar quark pair to W pair diagramhexagon-box diagramgluon fusion to Z gamma diagram
the equal-mass kite diagramcrossed three-loop light-by-light diagramthe Calabi-Yau banana diagrammulti-loop ladder diagram

Related projects run on essentially the same technology: the four-point energy correlator, cosmological correlators — including, in closed form, the one-loop coefficient whose master integrals are elliptic and whose elliptic periods cancel — the anisotropic Watson integral of condensed-matter physics, and string theory's modular graph functions.

The strength of natural selection — The harm a new mutation does is read from how rare it stays across huge groups of people, through a sixty-year-old formula that could not be evaluated exactly where selection is strong and samples are large, and through software that had never been checked against an exact answer. BootLoops evaluated the formula exactly, in rational arithmetic, at any selection strength, and audited the standard programs against it. The science: a fit to roughly 730,000 human exomes in the gnomAD compendium, 1.46 million genome copies, places most gene-disabling mutations in the strongest-selection range the data can resolve — a rare-variant regime that classical sample sizes were too small to reach — and the standard packages prove mostly reliable at the sample sizes they were built for, with a map of where each loses accuracy.

A scattered population of about forty slender three-dimensional DNA double helices at random orientations, twenty-four base pairs each, color-coded A, T, G, C — every copy identical except one, haloed, where a single circled base pair differs: a rare variant among 1.46 million genome copies

An exact tie among the apes — Bayesian phylogenetics scores each candidate family tree by the probability of the DNA under that tree, averaged over branch lengths, a number estimated by Monte Carlo since the 1990s. With the evolutionary biologists Scott Edwards and Paul Lewis, BootLoops showed that for four or five species under the simplest models of DNA change it is an exact fraction, computable with Feynman-integral techniques, and released the computation as the JaCK & Jill package. The science: on one mitochondrial gene of human, chimpanzee, gorilla, orangutan and gibbon, all fifteen trees are scored exactly; the accepted tree and its image with human and gorilla exchanged give the same ratio of integers, a tie no sampler can establish, and the tree pairing human with gorilla, which led without the gibbon, ranks third.

Two five-species ape trees with animated silhouettes at the tips — a standing human who taps a foot and glances left and right, a squatting slate chimpanzee that scratches its head, a sitting gorilla that beats its chest twice, an orange orangutan hanging from its branch by one arm and swaying, and a buff gibbon swinging by one long arm; on the left the accepted tree (human with chimpanzee, then gorilla, orangutan, gibbon), on the right the same tree with human and gorilla exchanged, joined by a violet equals sign reading exactly tied, the same ratio of integers; underneath: one gene, all fifteen five-ape trees scored exactly, these two tie, pairing human with gorilla ranks third

Medicare's star ratings — Medicare Advantage pays bonuses by a star system whose rulebooks state an exact optimality criterion for the cut points, computed in practice by a heuristic. BootLoops solved the star-assignment criterion exactly — the optimum has been computable since 1958 — and measured the production output against it. The science: the published thresholds fail their own stated criterion in 104 of 123 recent computations, and stars assigned from the true minimum would reallocate roughly $0.9–1.4 billion in bonus payments across payment years 2025 through 2027 — 34 of 1,724 published contract-year ratings change bonus status, 16 up and 18 down. No one made an arithmetic error; the criterion and the heuristic simply part ways, invisibly, until an exact solver looks.

One peer group of 177 hospitals, each a dot at its standardized 2026 summary score, with CMS's shipped star bins as a grey ruler above the dots and the certified-optimal bins as a blue ruler below — every shipped cut point sits to the right of its certified-optimal counterpart
One peer group from the paper: 177 hospitals, each a dot at its 2026 summary score. The grey ruler is the five star bins CMS shipped; the blue ruler is the certified optimum. Every one of the year's twelve shipped cut points sits above its certified-optimal counterpart, and each hospital caught between a paired solid and dashed edge gets one star fewer than the optimum assigns — 213 hospitals in 2026 alone. (Figure 1 of the paper.)

A forest deciphered — On Barro Colorado Island in the Panama Canal, every tree in fifty hectares of forest has been counted eight times since 1982. Neutral theory, in which species differ only by luck, predicts the overall shape of rarity and abundance in such a forest but not which species gain trees and which lose them, and it underestimates the pace of change. BootLoops put the theory to an exact test on those censuses and then, with the ecologist James O’Dwyer, joined the two remedies proposed for its failings — environmental variation and the species’ measured life histories — with immigration in one minimal model. The science: the model predicts the dynamics and abundances of the island’s species; environmental variation enters as a single common effect on growth shared by all species, and regional commonness together with a simple competition–colonization trade-off predicts relative abundances. What the censuses require is much simpler than the model’s ingredients would allow; whether a comparably simple model works in other forests is open. A seven-minute film on the page tells the story.

A cartoon strip of tropical forest — stylized trees and palms in purple, slate, orange and black on a lilac ground line, a sloth hanging from a branch, a toucan perched on a cypress, two birds overhead

How Earth’s air got its oxygen — For the first half of the planet’s history the air held almost no oxygen, and the rise that began about 2.4 billion years ago, the Great Oxidation, is usually drawn as a single permanent step. With the geochemist David Johnston, BootLoops assembled what different groups had measured separately — the photochemistry, the climate thresholds, the dated rocks — into one minimal ocean–atmosphere model whose every input is fixed by published measurements, the present-day steady state, the dated onset of the transition and the glacial record. The science: in the model the rise is not a step but an oscillation entangled with four glaciations from 2.45 billion years ago, each returning the air to anoxia if oxygen made beneath the ice is consumed there; oxygen becomes permanent only at the last deglaciation, 2.22 billion years ago. The model’s third glaciation ends about 20 million years later than dated ash beds allow, and carbon dioxide levels computed from two ancient soils inside the transition are consistent with it. A short film on the page tells the story and walks through the equations.

A cartoon of the transition: on the left an orange, hazy early Earth with a smoking volcano under a pale sun; in the middle alternating bands of orange and blue sky above an ice sheet, the oscillation paced by glaciations; on the right a blue sky with clouds over open ocean

Other applications — The six above are a sample. The same instruments run through cosmology’s large-scale structure, where the two-loop power spectrum of galaxy clustering is computed and tabulated once for every cosmology; black-hole scattering and the memory gravitational waves leave in spacetime; the demography of sunspots; genes that fire in bursts in single cells; mixture models in statistics; shared authorship — the Federalist Papers, the Shakespeare canon, the book of Isaiah — and the Voynich Manuscript; language history and typology; string theory’s flux vacua and string integrals; clinical-trial designs; and America’s ozone limit. In economics, an LLM workflow reproduces, improves, and extends published research from replication packages. Each has its own page under Web summaries.

Acknowledgements

The science produced by BootLoops would not have been anywhere near as interesting without guidance and collaboration from: Isaiah Andrews, Nima Arkani-Hamed, Michael Desai, Scott Edwards, Noam Elkies, Cecilia Garraffo, Matthew Gordon, Thomas Grimm, Martin Hemberg, Mikhail Ivanov, David Johnston, Gary King, Paul Lewis, Brendan Meade, Amara McCune, Joe Pater, Subhabrata Sen, Siddharth Mishra-Sharma, James O'Dwyer, Kevin Ryan, Jesse Shapiro, and Xiaoyuan Zhang.

Thanks also to Anthropic for providing the tokens and compute that powered this project. Results and views are not endorsed by Anthropic.