Interpretability, the set of tools that peek inside how models reach answers, used to live in academic papers. Goodfire just raised $150M to sell it as a $1,000/month product.
Goodfire is asking AI labs to pay $1,000 a month for a tool that tries to read what a model is "thinking" as it processes data. The product, called Silico, packages a specific, contestable scientific claim into a paid subscription: that post-training mostly amplifies what a model already knows rather than installing new skills, and that the right kind of inspection can predict which training preferences will backfire before a single GPU-hour is spent.
The claim comes from a small but fast-growing research program at Goodfire, the San Francisco interpretability lab co-founded by Dan Balsam. On the latest Cognitive Revolution podcast, Balsam described the underlying mechanism, what Goodfire calls Predictive Data Debugging, or PDD, as a way to look at a preference dataset (the labeled comparisons used in RLHF, the reinforcement-learning stage that turns a base model into a chatbot) and predict which behaviors the training loop will amplify or suppress without running the training at all. The system then traces those predictions back to the specific training examples responsible.
If post-training really does mostly tweak the relative likelihood of behaviors the model can already produce, then the data fed to it is the lever, not the algorithm. That is the load-bearing claim behind Goodfire's Predictive Data Debugging work, and the paper's authors, Leon Bergen, Usha Bhalla, and Max Loeffler, argue it is empirically true. They describe the relationship between data filtering and reward shaping as a deep isoperativity: cut a harmful preference from the dataset and you get a comparable effect to penalizing it in the reward function, with comparable off-target damage to other behaviors.
PDD grew out of an older line of work the field calls "Toy Models of Superposition," a roughly three-year-old framing that argues neural networks don't store concepts as one-hot switches but as sparse mixtures of overlapping subspaces, what Goodfire calls "concept manifolds." The practical consequence is that steering a model "off-manifold," or pushing it in a direction its internal geometry was not built for, is the most common way interventions fail. Goodfire says it has replicated PDD at the scale of Moonshot's Kimi k3 and Zhipu's GLM, two of the larger open-weight model families out of China, a non-trivial claim in a field still fighting the "toy-model science" reputation.
The product is what turns this from a research update into a market story. Silico, sold at $1,000 per month, is pitched as a long-horizon asynchronous research agent: a system that can spin up training jobs, run sparse-autoencoder and probe experiments, and orchestrate compute across different hardware fabrics while a human researcher sets the direction. Goodfire's pricing page describes it as an interface to the company's interpretability stack, including training, SAE and probe training, neural-geometry mapping, and compute orchestration, rather than a turnkey debugging tool.
A third-party writeup from AlphaSignal reports that Silico "cuts hallucinations by 37% in cited usage." That figure comes from a secondary outlet, not from a Goodfire benchmark, and is best read as a directional customer outcome rather than a reproducible result. Goodfire has not published a hallucination leaderboard.
The funding marker landed at the same moment. On February 5, 2026, Goodfire announced a $150 million Series B at a $1.25 billion valuation, the announcement-of-record for the round. AlphaSignal separately cites a $207 million figure; until the company or its investors reconcile the two, the $150M and $1.25B combination is the more reliable marker.
Whether the field-maturation claim holds is a separate question. The mechanism Goodfire is selling is interpretive, comes largely from Goodfire-affiliated sources, and is contested in the broader interpretability community, where competing research groups publish their own takes on what post-training actually does. Goodfire has commercial interest in the worldview being right, because Silico is priced against it. A $1,000-a-month product built on a contested theory is not a falsification, but it is a positioning bet.
What to watch is whether the bet pays off in customer logos rather than in explainer posts. Goodfire has not disclosed paying customers; the funding and the price point are the receipts, not the adoption curve. A measurable training outcome from a frontier lab or an enterprise team in the next two quarters would turn the field-maturation claim from a thesis into a market.