Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost
Goodfire is selling internal AI-agent monitors that it says are much cheaper than having another model reread every step.
The system uses small probes inside a model’s computations and escalates only flagged activity to a separate AI reviewer. In Goodfire’s Kimi K3 tests, about 1,500 monitored sessions cost roughly $51, versus $233 for a cheaper external checker and about $10,000 for a top-tier one. The company says the probes caught 94% of malicious hacking sessions, with 8.7% of harmless sessions sent for review. Baseten customers can choose risks such as offensive hacking, biosecurity misuse, and reward hacking, then set responses from logging to refusal. TechCrunch AI's note
The system uses small probes inside a model’s computations and escalates only flagged activity to a separate AI reviewer. In Goodfire’s Kimi K3 tests, about 1,500 monitored sessions cost roughly $51, versus $233 for a cheaper external checker and about $10,000 for a top-tier one. The company says the probes caught 94% of malicious hacking sessions, with 8.7% of harmless sessions sent for review. Baseten customers can choose risks such as offensive hacking, biosecurity misuse, and reward hacking, then set responses from logging to refusal. TechCrunch AI's note
score 5