Harvard's Matthew Schwartz argues researchers should stop fighting current AI and pick problems that match what the models actually do well.
The post opens with a diagnosis: a year of trying to make current AI useful in a physics research context, hitting the same wall most working researchers hit. The models were fluent and fast. The published wins were dramatic. But the day-to-day work of building a calculation, checking a derivation, or connecting a result to a neighboring field kept stalling. The post names the gap an "impedance mismatch" in a guest post on Anthropic's research blog. The phrase, borrowed from electrical engineering, marks the moment when two systems that should talk to each other do not.
The diagnosis matters because the usual response to a stalling tool is to push harder: longer prompts, more context, more retries. The post describes trying that, and it works, until it does not. The reframe is the inversion most frustrated researchers eventually arrive at. Instead of dragging the model to a hard problem, change the problem until it matches what the model actually does well. The post calls the result a "Claude-shaped problem." The brand tie is awkward, but the diagnosis generalizes. Current large language models (LLMs) are statistical systems trained on text, and they tend to excel at tasks with a tight feedback loop: a well-defined input, an exact, checkable output, and a path the model can grind through and verify. Pure mathematics fits that shape. Most experimental science does not, because the inputs are messy, the answers are graded by domain experts rather than by a proof checker, and the interesting questions are not yet well-posed.
The post is explicit about the limit. As it puts it: "Claude and GPT are good at science, but they are not scientists." The distinction shows up in the literature. Most documented LLM wins in research are mathematical: AlphaProof at the 2024 International Mathematical Olympiad, Lean-based theorem provers cracking undergraduate problems, AI-coauthored results in combinatorial theory. The wins are real and the tools are useful, but the bottleneck in the lab is rarely the part of a problem that a proof checker can grade.
The response was to build a small toolkit that fits the model. BootLoops, open-sourced under the group, is a thin Python layer that turns the kind of exact, calculation-heavy work a theoretical physicist actually does into a structured loop the model can drive. A companion site hosts the examples. The artifact is modest: a few hundred lines of glue around the model's code execution, with explicit checkpoints where a human or another program can verify the result. The point is to give the model a problem it can finish and to keep a human in the loop at every step that matters.
The interesting move is what happened after the toolkit was working. The post reports that when BootLoops was pointed at problems outside physics, the model turned up plausible connections to ecology, population genetics, materials science, and roughly a dozen other fields. The post then notes the raw hits were "often technically correct but scientifically unremarkable at first." This is the load-bearing caveat. Cross-domain suggestions from a language model are easy to generate and easy to overrate. The post's own recovery is to bring a domain expert into the loop, someone who can tell which connection is a real question and which is a rephrasing of common knowledge. The steer is where the science happens.
That is the part working researchers should take from the post, more than the toolkit itself. The recipe has three steps: pick a problem whose structure matches the model's strengths, build a thin scaffolding that lets the model drive the loop, and refuse to ship a result that a domain expert has not vetted. The first step rules out most of what an experimentalist does. The second step is the engineering. The third is where the science still belongs to a human.
A Hacker News thread on the post drew about a hundred comments. Some readers flagged the framing as "AI psychosis." Others pointed to the BootLoops repository as the part worth taking seriously. The verdict from the field is still out: BootLoops is a single group's tool, and there is no public track record yet of outside labs using it on their own problems. Treat the playbook as a working hypothesis from one researcher, not a settled method.
AI in science coverage tends to fixate on the math stunts: the olympiad wins, the Lean proofs, the AI co-authors. Those wins are real. They also leave most working labs cold, because the impedance mismatch is not a model-quality problem, it is a problem-shape problem. The post's contribution is the reframe, and the artifact that came out of it. Other labs will judge the playbook by whether they can run BootLoops on their own problems and get something back worth keeping.