FIELD NOTES / AI
01 SEPT 2026 / REWARD-HACKING PROBES02 AUG 2026 / $1M IN RESEARCH GRANTS03 FEB 2026 / $150M SERIES B

Company profile / AI interpretability

The Most Useful Thing an AI Model Knows May Be What It Cannot Say

Goodfire wants to open the black box of AI. At Rakuten, that meant a cheaper privacy guardrail; in Alzheimer’s research, it meant a short list of biomarkers worth testing.

Imagine you have hired an unusually clever assistant. It can spot a disease in a blood sample or answer a customer at midnight. Ask it why it made a particular choice, however, and you may get a plausible story rather than a reliable account of what happened inside. That is the awkward bargain at the center of modern AI: extraordinary performance, scant mechanical understanding. Goodfire, a San Francisco research company founded in 2024, has made a business of examining the machinery.

Its field is called mechanistic interpretability. The name sounds as though it ought to come with a lab coat; the ambition is closer to opening a watch. Goodfire studies the internal signals of neural networks, tests what those signals mean, and tries to use them to change or monitor behavior. Eric Ho, the chief executive, founded the company with Daniel Balsam, its chief technology officer, and Thomas McGrath, its chief scientist. Ho and Balsam had worked together at RippleMatch; McGrath helped start the interpretability team at Google DeepMind.

In brief

  • Goodfire sells tools and research support for reading and shaping AI model internals.
  • Rakuten used its Ember platform to build a privacy filter for an AI agent platform.
  • Prima Mente used it to turn an Alzheimer’s model's opaque predictions into candidate biomarkers.
  • Silico, its newer research agent, runs experiments and produces reports a researcher can inspect.

The embarrassing question after a good answer

Most teams can measure a model's output. They can collect bad answers, change prompts, fine-tune on more examples, and measure again. Yet the question that follows a failure remains stubborn: what inside the model caused this? Goodfire's tools look for interpretable patterns in the model's activations, the internal computations produced while it works. Researchers can then train small detectors, often called probes, to recognize a pattern associated with a behavior. They can also test whether changing that pattern changes the output.

That is a different vantage point from a dashboard of prompts and responses. An output monitor sees what the model said. An internal probe may catch a signal before the next action, provided the operator has access to the model's internals and has trained and tested the probe for the task. Neither tool abolishes uncertainty. The point is to gain another instrument, especially when reading every agent action or asking another large model to judge it would be too slow or costly.

A privacy filter with a peculiar training diet

Rakuten gave Goodfire a particularly useful test. It wanted to catch names, addresses, phone numbers and email addresses before they passed from its AI agent platform to downstream model providers. The detector had to be fast enough for live use. It also had to be trained on synthetic data, because real customer data could not be used for that purpose. Those conditions tend to expose a fragile method quickly: a filter that shines on tidy fabricated examples can miss noisy, multilingual production traffic.

The two teams tried several methods with Goodfire's Ember platform. The one that held up best when moving from synthetic examples to real data used sparse autoencoder probes. A sparse autoencoder is a way to break a model's busy internal activity into more legible features; a probe checks whether the feature associated with a target behavior is present. Rakuten deployed the resulting guardrail on its agent platform. In Goodfire's account, the probes had comparable performance to a large language model used as a judge while costing 15 to 500 times less in the tested setups.

44M+Rakuten monthly active users in Japan
15-500×Reported cost saving against comparable LLM judges
2024Partnership began
Goodfire chart comparing the cost of PII detection approaches in its Rakuten project
The unromantic detail that makes this interesting: a guardrail has to be cheap enough to stay switched on. Goodfire's published cost comparison for the Rakuten project.

The comparison should be read with its conditions attached. The reported savings are for Rakuten's task and the methods Goodfire tested, not a universal price list for every model. But the general lesson is easy to borrow: define the expensive failure, establish the baseline, train under the data constraints you actually have, and test on messier traffic than the training set. A probe earns its place only when it beats a simpler option on the measures that matter in production.

When the model has found something its makers have not

The second Goodfire story begins with a successful model and an unsatisfied scientist. Prima Mente had trained an epigenomics foundation model for early Alzheimer’s detection. It performed well, yet its prediction did not tell the team which biological signals deserved precious time in a wet lab. A classifier can point toward disease without explaining whether it has found a meaningful marker, a confounding shortcut, or a mixture of both.

Goodfire researchers embedded with Prima Mente's team after model training. They trained sparse autoencoders, traced predictions back to interpretable signals, tested whether those signals generalized to new patients, and removed dominant signals in experiments to reveal subtler ones. The work produced a candidate class of blood-borne biomarkers. They are undergoing experimental validation; they are not yet a clinical test. That distinction matters. The model supplied hypotheses, and the lab still has to decide whether nature agrees.

“Nobody understands the mechanisms by which AI models fail, so no one knows how to fix them.”Eric Ho, Goodfire co-founder and CEO

Goodfire has also worked with Arc Institute on Evo 2, a genomic foundation model, announced a genomic medicine collaboration with Mayo Clinic, and partnered with Radical AI on materials discovery. These are not identical customers. What joins them is the possibility that a model trained on complex scientific data has learned a useful relationship that its headline prediction cannot reveal. Interpretability becomes a way to ask a better experimental question.

Goodfire team gathered for a group portrait
Goodfire's team. A company devoted to reading artificial minds still needs a room full of human ones.

From an expert service to a research desk

Goodfire's first visible commercial shape was Ember, the interpretability platform used in those customer projects. Its newer product, Silico, packages more of the research process. A user starts with a question; the system proposes a plan, runs an experiment, records its steps on a timeline, and produces a report with evidence to inspect and share. The documents show applications from protein structures to robot control and model safety. Researchers still have to judge the question, the controls and the conclusion. Silico's promise is to make the labor between them less forbidding.

The prices reveal the intended buyers. Eligible academic, nonprofit and safety labs can buy Silico for $1,000 a month. Enterprise teams get custom pricing, dedicated researcher support and organization controls. Goodfire also does close research work with partners. The company reports no public revenue figure or customer count, so the size of that business is an open question. Its funding is less mysterious: a $50 million Series A in 2025 was followed by a $150 million Series B led by B Capital in February 2026, at a $1.25 billion valuation.

The lab publishes research, maintains open repositories and has offered up to $1 million in free Silico usage for selected academic and nonprofit researchers. That is partly a research culture and partly a route into a small, specialized market. Goodfire needs researchers to improve interpretability methods, and customers need those methods to survive the untidy conditions outside a paper. The loop can help both sides, if the experiments remain honest.

The useful limit of looking inside

A September 2026 Goodfire study offers a neat illustration of both promise and restraint. Its researchers built probes for internal signals associated with agents gaming their rewards. On some tests, the probes caught cheating missed by monitors reading the model's written reasoning; on another comparison, they caught less. Goodfire describes a combined arrangement in which cheap probes screen widely and expensive language model reviews handle suspicious cases. There is no single magic microscope here. The instrument has to be calibrated against the model, task and false alarm rate.

That caveat is also a practical map for anyone tempted to copy the idea. Internal monitoring needs access to activations; a closed model available only through an API may not offer it. A probe needs data that represents the behavior of interest and tests on unfamiliar cases. A biological insight needs independent laboratory confirmation. When those conditions hold, the extra view can be valuable: a production filter can be faster, a scientist can choose better experiments, and a model builder can investigate a failure instead of merely collecting another bad answer.

Goodfire's wager is that AI will eventually be engineered with more of the control we expect from software. For now, its most persuasive work is narrower and more interesting. It takes a model that appears to know something, then asks whether that knowledge can be located, tested and put to use. A black box is tolerable until the day you need to repair it. Goodfire is betting that day has already arrived.