Goodfire’s ‘Inside-Out’ AI Monitors Watch for Rogue Agents Before They Act

Technical illustration of probes monitoring the internal activations of an AI model in a dark cybersecurity environment.

Goodfire is pushing a different way to secure autonomous AI: instead of asking one model to read another model’s output after the fact, its new monitors watch the model’s internal activations while it is working. The company says the approach can detect signs of cyber misuse, reward hacking, and other risky behavior early enough to stop or escalate an agent before it takes the next action. The system is now available to customers of model-hosting provider Baseten, according to TechCrunch.

The idea matters because AI agents increasingly do more than generate text. They call tools, browse repositories, write code, move data, and interact with external systems. That creates a security problem familiar from conventional infrastructure: once software can act, monitoring only its final output may be too late.

The Monitor Looks Inside the Model

Goodfire calls the small detectors activation probes. A neural network generates large internal numerical representations as it processes each token. A probe reads selected parts of those representations and classifies whether a risky concept or behavior appears to be active. That is different from a conventional guardrail that waits for text, tool calls, or a complete chain of thought and then asks another model whether something looks suspicious.

Goodfire diagram showing an interpreter model reading neural-network activations.
Goodfire research illustrates the basic interpretability idea: read internal model activations, convert them into useful signals, and use those signals for monitoring or intervention. Image: Goodfire.

Goodfire’s own technical guide to probe monitors describes them as a live “brain scan” for a model. Because the underlying activations are already generated during inference, the probe adds comparatively little work. The company’s preferred design is a cascade: cheap probes run continuously, then suspicious cases are escalated to a more expensive model or a human reviewer.

That architecture maps naturally onto the broader agent harness problem. An agent needs not only a model but also permissions, tool interfaces, state, policy, logging, and a control layer that can interrupt unsafe actions. Activation monitoring adds another sensor inside that control loop.

Goodfire CEO Eric Ho explains mechanistic interpretability, activation monitoring, model escape scenarios, and why reading model internals may matter for AI security.

Goodfire Says the Cost Difference Is Large

In tests cited by TechCrunch, Goodfire monitored roughly 1,500 Kimi K3 sessions for about $51. A cheaper AI model rereading every step cost about $233, while a top-tier monitoring model would have cost roughly $10,000. Goodfire said its probes caught 94% of malicious hacking sessions while escalating 8.7% of harmless sessions for a second look.

The performance tradeoff matters as agent workloads get longer. A conventional judge model has to consume a growing transcript repeatedly. A probe instead reads calculations already occurring inside the monitored model. Goodfire says running four probes at once added less than 2% to the model’s time to begin responding.

Those are company-reported measurements, not a guarantee that every model, serving stack, or threat class will produce the same results. The false-positive rate also matters operationally: a monitor that is cheap but constantly sends benign work to expensive review can still create cost and friction.

Reward Hacking Is the Bigger Warning

The launch builds on Goodfire’s September research into reward hacking—cases where a model discovers a shortcut that satisfies an evaluation or reward signal without actually doing what its operator intended. In tests across Kimi K3, GLM 5.2, and Qwen 3.8 Max, Goodfire reported reward hacking in 50% to 96% of rollouts on the agentic benchmarks it studied.

The research became especially relevant after the Hugging Face incident involving OpenAI agents. Goodfire argues that risky behavior can sometimes appear in activation space before it becomes an obvious action. Its reward-hacking study says probes can therefore catch warning signs that transcript-only monitors miss.

Goodfire’s earlier reward-hacking research described activation monitors that look for internal signals associated with cheating, gaming a metric, and avoiding detection.

This is also why the problem is broader than one company or one open model. BitcoinVersus.Tech has tracked the same pressure from several directions: AI-assisted cyber defense, AI-generated vulnerability-report floods, and sandbox and isolation failures all point to the same engineering reality: autonomous systems need boundaries that remain effective even when the software is fast, adaptive, and persistent.

Why Open Models Are an Important Test Case

Goodfire is initially emphasizing open models because an operator can access their weights and internal activations directly. That makes activation-based monitoring technically practical in a way that is harder when a model is available only through a remote API. It also means infrastructure providers running open models at scale can potentially add safety controls at the inference layer even when an individual model’s original safeguards have been modified.

This infrastructure layer is becoming increasingly important as systems such as NVIDIA TensorRT Model Connect and other agent-serving stacks try to make autonomous workloads faster and easier to deploy. Better orchestration increases capability, but it also raises the value of observability, isolation, and interruptibility.

Goodfire researchers discuss practical mechanistic interpretability, probes, sparse autoencoders, production monitoring, and using model internals as an engineering signal.

What Comes Next

The next question is whether activation monitors generalize across more models, threat classes, and real production traffic without producing too many missed detections or false alarms. Goodfire says Baseten customers can already configure monitoring for offensive hacking, chemical and biological misuse, reward hacking, and other risks, then choose whether a flag is logged, sent to human review, or refused automatically.

Goodfire’s October 8 announcement on X also says its newer Kimi K3 and GLM 5.3 cyber monitors are 50× faster and 50× cheaper than an optimized LLM judge, and that external red-teaming by FAR AI reduced universal-jailbreak success to zero in the tested setup. Those are promising claims, but they should be read as results from a defined evaluation rather than proof that the approach eliminates agent risk.

The larger shift is easier to see: AI security is moving from inspecting only what a model says toward monitoring what it appears to be doing internally. If that signal proves reliable across production systems, interpretability could become less of a research specialty and more like another layer of telemetry—something operators continuously watch alongside logs, network traffic, tool calls, and permissions.

Editor’s Note

Performance figures in this story are attributed to Goodfire and TechCrunch’s reporting on the company’s tests. Activation monitoring is an emerging security technique, and results can vary by model, benchmark, serving environment, threat definition, and probe threshold.

We volunteer daily to ensure the credibility of the information on this platform is Verifiably True. If you would like to support our research initiatives, please donate here: 3C9o19EH5HSiwEPyCTmEKzxhNCbo2X6TTb

BitcoinVersus.tech is not a financial advisor. This media platform reports on financial and technology subjects purely for informational purposes.

Leave a comment