Goodfire released Silico on April 30, 2026, a tool that lets engineers inspect and adjust the parameters of a large language model while it is being trained. The San Francisco startup is pitching it as the first off-the-shelf product that covers every stage of model development, from dataset construction to post-training debugging. Goodfire is competing directly with the in-house interpretability teams at Anthropic, OpenAI, and Google DeepMind — the three frontier labs that have been driving the field.
The pitch is engineering discipline for an industry that still ships models nobody fully understands. Goodfire CEO Eric Ho told MIT Technology Review the gap between how widely models are deployed and how well they are understood keeps widening. "I think the dominant feeling in every single major frontier lab today is that you just need more scale, more compute, more data, and then you get AGI and nothing else matters," Ho said. "And we're saying no, there's a better way."
Silico sits on top of mechanistic interpretability, the technique of mapping individual neurons inside a model and tracing the pathways between them. Until now, that work has largely been an internal research function at the biggest labs. Goodfire is packaging the tooling and selling it, with pricing set case by case rather than published.
Key facts
- 01Goodfire released Silico on April 30, 2026, billing it as the first off-the-shelf mechanistic interpretability tool spanning data, training, and debugging.
- 02In a Goodfire test, boosting transparency-linked neurons flipped a model's answer 9 out of 10 times on whether to disclose deception affecting 0.3% of 200 million users.
- 03The tool competes with in-house interpretability work at Anthropic, OpenAI, and Google DeepMind, the three frontier labs leading the field.
- 04Goodfire surfaced a neuron in the open-source Qwen 3 model that, when activated, reframes outputs as explicit moral dilemmas.
- 05CEO Eric Ho says agents now do much of the interpretability labor that previously required human researchers, making the platform commercially viable.
The product leans heavily on agents to automate the grunt work. "Agents are now strong enough to do a lot of the interpretability work that we were doing using humans," Ho said. "That was kind of the gap that needed to be bridged before this was actually a viable platform that customers could use themselves." That shift is what turns interpretability from a research luxury into something a smaller engineering team can actually run.
“Boosting neurons tied to transparency flipped a model's answer on disclosing deceptive behavior nine out of 10 times, even when the case involved 0.3% of 200 million users.”— Jaeden Schafer
The most striking demo Goodfire released involves a model asked whether a company should disclose that its AI behaves deceptively in 0.3% of cases affecting 200 million users. The model said no, citing commercial risk. After Goodfire researchers identified neurons associated with transparency and boosted their activation, the answer flipped from no to yes nine out of 10 times. "The model already had the ethical reasoning circuitry, but it was being outweighed by the commercial risk assessment," Ho said.
Other examples are weirder. Inside the open-source Qwen 3, Goodfire found a single neuron tied to the trolley problem; activating it pushed the model to reframe its outputs as moral dilemmas. "When this neuron's active, all sorts of weird things happen," Ho said. The same approach can explain known model failures — for instance, why many LLMs insist 9.11 is greater than 9.9. The culprit may be neurons associated with Bible verses, where 9.9 precedes 9.11, or with versioned code repositories. Filter that signal out at training time and the math improves.
Silico is aimed at the tier of companies that cannot afford to staff a dedicated interpretability lab. Leonard Bereska, a mechanistic interpretability researcher at the University of Amsterdam, said the tool looks useful and could matter for safety-critical work in healthcare and finance. "Silico arms the next tier of companies, where the value is not having to hire interpretability researchers," he said. "Frontier labs already have internal interpretability teams."
Bereska is less sold on Goodfire's framing. "In reality, they are adding precision to the alchemy," he said. "Calling it engineering makes it sound more principled than it is." That is the central tension. Silico exposes more knobs than developers had before, but it does not yet promise the kind of guarantees that real engineering disciplines deliver. Adjusting a transparency neuron is not the same as proving a model will behave a specific way in production.
There is also a commercial question Goodfire has not answered. The company declined to share pricing, and the bespoke quoting model suggests Silico is being sold to a small number of well-funded customers rather than as a self-serve product. The frontier labs Goodfire wants to compete with are spending billions on training runs; whether mid-tier model builders will pay enough to support an interpretability platform at scale is unproven.
If Silico works as advertised, the more interesting consequence is structural. Ho's bet is that interpretability tooling lowers the barrier to building credible custom models, the same way better compilers and debuggers expanded who could ship software. "If we can make training models a lot more like building software, there's no reason why there can't be many more companies designing models that fit their needs," he said. That cuts against the prevailing scale-is-all-you-need narrative coming out of the frontier labs, and it is the only version of the AI market in which a startup like Goodfire ends up mattering.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




