OpenInterpretability
  • Ekbasis
  • Guide
  • Pricing
  • Research
  • Lab
  • Tools
  • Notes
  • Registry
  • Manifesto
Sign inGet API keyAPI key
OpenInterpretability

Ekbasis API: know what an action will do before your agent runs it. From an independent lab for AI-agent safety. Open weights, Apache-2.0.

Get your API key
Ekbasis API
  • Console · sign in
  • Getting started
  • Pricing
  • For agents (agents.md)
  • Results and models
  • Client and CLI
  • Cookbook
The lab
  • About the lab
  • Open tools
  • Observatory
  • ProbeBench
  • InterpScore
  • Academy
Research
  • Manifesto
  • Roadmap
  • Papers & posts
  • Docs
Community
  • GitHub org
  • Notebooks
  • SDK (openinterp)
  • HuggingFace
  • Twitter / X
© 2026 OpenInterpretability — Apache-2.0 for code, CC-BY 4.0 for docs.Built in public.
Services live · SaaS waitlist open · Q3 2026

Six Diagnostics for autonomous crypto agents.

Your LLM agent moves capital. Its probes, monitors, and steering interventions need to survive their own diagnostics — or they're shipping a false sense of safety.

We built the protocol from 11 studies on Qwen3.6-27B — six walked back. We now audit yours.

Apply for audit — $5KJoin SaaS waitlist (Q3 2026)Read the Six Diagnostics →
Four failure modes — already documented

Your safety stack is shipping at least one of these.

We caught all four in our own work on one model. Six of our claims walked back. The protocol is what survived.

Probe-detected-then-not

env_coupling

Probe AUROC 0.86 in dev collapses to 0.30 in production. Env coupling between training env and live inference env breaks the direction.

Steering looks causal but is softmax shift

saturation_direction

Direction "flips" outputs at α=+5; Δrel after control-token normalization is −0.046. Uniform temperature change, not lever.

CoT-redirect causes obfuscation

template_lock (output-CoT decoupling)

Intervention clears 88% of verbalized sabotage but the action still ships sabotaged. Monitor goes blind.

Top-k probe hits 0.85 recall from marginal

marginal_fit

Predictive SAE recall@1024 = 0.85 reproduced ±0.03 by shuffled-source baseline. Probe learned marginal, not signal.

Full registry of walked-back claims: openinterp registry/failed-replications →

Live offer · $5,000 · 1–2 days turnaround

Manual Six Diagnostics audit of your agent stack.

You send us your probes, monitors, and any steering interventions currently in your agent's pipeline. We run the Six Diagnostics protocol on each. You receive:

  • A signed causal_report.json per the OpenInterp Eval Standard v0.1
  • A probe_card.json documenting every probe/monitor in your pipeline + the baselines it passed
  • A written recommendation memo (PDF, ~10 pages) covering the failure modes detected and how to mitigate
  • Optional: an intervention_trace.json record of any in-loop steering you currently run
  • Apache-2.0 reproducer scripts so you can re-run the diagnostics yourself
Price
$5,000
Turnaround
1–2 days
License
Apache-2.0 reproducers
NDA
On request
Apply for audit
Coming Q3 2026 · waitlist open

AgentGuard SaaS — hosted Six Diagnostics endpoint.

Webhook-callable safety gate. Your agent POSTs a candidate action; you get back PROCEED / FLAG / BLOCK with a causal_report.json reason. Designed for autonomous on-chain agents managing capital. First customers recruited from the audit cohort.

Free
$0
10 checks/day · evaluation only
Production
$99/mo
10K checks · 99.5% SLA · webhooks
Unlimited
$999/mo
unlimited · custom probes · dedicated
Join waitlist
Built by

The OpenInterp team — 11 studies on Qwen3.6-27B since launching OpenInterp in April 2026, 6 honest-negative walk-backs, shipping artifacts: FabricationGuard, agent-probe-guard, ProbeBench, and the OpenInterp Eval Standard. Apache-2.0 throughout.

FabricationGuard · AUROC 0.88 cross-taskagent-probe-guard SDK · detect-only by designFailed-Replication Registry · 6 entries

When the capital is real, the diagnostics should be too.

We're recruiting the first cohort of audit customers in 2026 Q2. Apply now and you also get priority access to the Q3 SaaS beta.

Apply for audit Browse Eval Standard