Your LLM agent moves capital. Its probes, monitors, and steering interventions need to survive their own diagnostics — or they're shipping a false sense of safety.
We built the protocol from 11 studies on Qwen3.6-27B — six walked back. We now audit yours.
We caught all four in our own work on one model. Six of our claims walked back. The protocol is what survived.
Probe AUROC 0.86 in dev collapses to 0.30 in production. Env coupling between training env and live inference env breaks the direction.
Direction "flips" outputs at α=+5; Δrel after control-token normalization is −0.046. Uniform temperature change, not lever.
Intervention clears 88% of verbalized sabotage but the action still ships sabotaged. Monitor goes blind.
Predictive SAE recall@1024 = 0.85 reproduced ±0.03 by shuffled-source baseline. Probe learned marginal, not signal.
Full registry of walked-back claims: openinterp registry/failed-replications →
You send us your probes, monitors, and any steering interventions currently in your agent's pipeline. We run the Six Diagnostics protocol on each. You receive:
Webhook-callable safety gate. Your agent POSTs a candidate action; you get back PROCEED / FLAG / BLOCK with a causal_report.json reason. Designed for autonomous on-chain agents managing capital. First customers recruited from the audit cohort.
The OpenInterp team — 11 studies on Qwen3.6-27B since launching OpenInterp in April 2026, 6 honest-negative walk-backs, shipping artifacts: FabricationGuard, agent-probe-guard, ProbeBench, and the OpenInterp Eval Standard. Apache-2.0 throughout.
We're recruiting the first cohort of audit customers in 2026 Q2. Apply now and you also get priority access to the Q3 SaaS beta.