BootLoops Toolkit Enables Claude-Shaped Scientific Computations Across Disciplines

TL;DR

Matthew Schwartz released BootLoops, an open‑source harness that lets Claude (and other LLMs) perform exact, large‑scale calculations across many scientific fields, turning Claude‑shaped problems into publishable results.


What is BootLoops?

BootLoops is a software harness that connects Claude’s code‑generation, mathematical, and data‑parsing capabilities to quantitative scientific workflows. It functions like Claude Code or Claude Science but is model‑agnostic, allowing any LLM to be used. The toolkit automates the entire pipeline: importing literature, porting algorithms to a common framework, generating code, running high‑precision numerics, and validating results against analytical or numerical benchmarks.

"BootLoops functions as a kind of harness for the LLM, much like Claude Code or Claude Science is a harness for Claude, or Codex is a harness for GPT." – Matthew Schwartz

Why BootLoops matters

Current LLMs excel at breadth of knowledge and coding but struggle with the impedance mismatch between what scientists need (conceptual insight, novelty) and what LLMs provide (exact calculations, data handling). BootLoops narrows this gap by focusing on Claude‑shaped problems—tasks that match the strengths of today’s agentic AI: massive knowledge, fast coding, and reproducible numerics.

Core Technical Approach

  1. Unified Framework – All imported methods (e.g., scattering‑amplitude calculations, Bayesian evidence integrals) are translated into a single codebase (Python/Julia) that Claude can extend.
  2. Semi‑Numerical Bootstrap – The toolkit implements the semi‑numerical bootstrap, which combines analytic constraints with high‑precision numeric evaluations (often >1,000 digits) to solve integrals that are otherwise intractable.
  3. Agentic Session Management – Each research project runs in a dedicated Claude Code session on Google Cloud VMs, with sub‑agents handling computation, file organization, and result validation. A master session orchestrates resource allocation and monitors for classifier‑triggered shutdowns.
  4. Failure‑Mode Guardrails – Protocols enforce strict success criteria (no unproven lemmas, explicit ETA testing, mandatory plot inspection) to mitigate Claude’s tendency to declare premature completion.

Highlighted Scientific Applications

Field Problem Type BootLoops Contribution Outcome
Mathematical Physics Elliptic Feynman integrals Ported semi‑numerical bootstrap, generated code for elliptic families 30 integrals solved (15 known, 15 novel) in weeks
Ecology Neutral biodiversity theory (Barro Colorado Island) Solved Etienne’s 2005 equation, compared to forest data Demonstrated species turnover 4.5× faster than neutral prediction; led to refined predictive model with expert James O’Dwyer
Population Genetics Selection‑shaped rare‑mutation integral Imported physics methods, applied to gnomAD dataset Validated technical result; later extended to 5.7 B mutation pairs, revealing gene‑conversion evidence
Economics Replication‑package validation Converted 4,452 papers’ code from MATLAB/Stata to open‑source, cross‑checked numbers Results documented in NBER working paper (W35782)
Linguistics Word‑stress database construction Scraped 6,072 languages, quoted source passages, built AccStack Public database with 160 k bibliographic entries
Phylogenetics Bayesian evidence for tree selection Accelerated exact evidence computation, identified ties within sampling error
Earth Science Great Oxidation Event modeling Integrated atmospheric, geochemical, climate equations Quantitative model spanning four glaciations
Sunspot & Starspot Demography Lifecycle analysis using demographic methods Applied human‑population techniques to solar data New characterization useful for exoplanet studies
Cosmology Large‑scale‑structure parameter fitting Implemented full two‑loop power spectrum and one‑loop trispectrum toolkit
Statistics Bayesian evidence for mixture models Derived precise approximations, enabling practical model selection

Workflow with Domain Experts

Schwartz emphasizes that Claude often produces technically correct but scientifically modest results. Expert collaboration refines these outputs into meaningful discoveries. Examples include:

  • James O’Dwyer (Plant Biology) helped reshape neutral‑theory results into a predictive life‑history model.
  • Michael Desai (Genetics) guided the shift from rare‑mutation integrals to gene‑conversion analysis.
  • Economists & Linguists supplied domain‑specific validation and contextualization.

Practical Lessons for Using Agentic AI

  • Define Success Rigorously – Require explicit criteria (e.g., no unproven lemmas, reproducible plots).
  • Verify Independently – Always inspect generated figures and run separate checks.
  • Guide Towards Novelty – Instruct Claude to prioritize newly possible problems rather than well‑trodden calculations.
  • Iterate with Experts – Use the AI to generate drafts, then let specialists steer the scientific narrative.
  • Manage Compute – Sessions are token‑intensive; allocate cloud resources wisely and monitor for classifier interruptions.

Outlook and Implications

BootLoops demonstrates that when the problem space aligns with LLM strengths, the bulk of technical work—coding, high‑precision numerics, data curation—can be automated. This frees researchers to focus on conceptual framing, hypothesis generation, and interdisciplinary synthesis. The approach also highlights a new credit model: human scientists provide the direction and validation, while the AI supplies the execution.

Future directions include expanding the harness to more model families, integrating automated literature‑review pipelines, and building community contributions via the open‑source GitHub repository.

Resources


This post is a guest contribution by Matthew Schwartz, a visiting researcher at Anthropic. BootLoops is not an Anthropic project.

Sources