OpenInterpretability
  • Ekbasis
  • Guide
  • Pricing
  • Research
  • Lab
  • Tools
  • Notes
  • Registry
  • Manifesto
Sign inGet API keyAPI key
OpenInterpretability

Ekbasis API: know what an action will do before your agent runs it. From an independent lab for AI-agent safety. Open weights, Apache-2.0.

Get your API key
Ekbasis API
  • Console · sign in
  • Getting started
  • Pricing
  • For agents (agents.md)
  • Results and models
  • Client and CLI
  • Cookbook
The lab
  • About the lab
  • Open tools
  • Observatory
  • ProbeBench
  • InterpScore
  • Academy
Research
  • Manifesto
  • Roadmap
  • Papers & posts
  • Docs
Community
  • GitHub org
  • Notebooks
  • SDK (openinterp)
  • HuggingFace
  • Twitter / X
© 2026 OpenInterpretability — Apache-2.0 for code, CC-BY 4.0 for docs.Built in public.
Independent AI-agent safety lab · pre-registered · honest walk-backs · permanent DOIs

Understand and control AI agents.

An independent lab studying how LLM agents fail on long-horizon tasks — and building the open tools that catch those failures before they do harm.

We publish positives and nulls with the same rigor: pre-registered, every number recomputed from public data, permanent Zenodo DOIs. The WANDERING arc explains how agents fail; AgentGuard and Ekbasis stop the failures that matter.

Read the research arcThe agent safety stack28 papers · permanent DOIsCode & data →
New · open modelApache-2.0 · weights, data and code

Ekbasis ἔκβασις

An open world model for agents

Before an action runs, it predicts what the action will do — will this lose work? will it fail? — with a calibrated probability, in one forward pass. Trained on what actions actually did.

Meet Ekbasis

Not a model that thinks. Not a model that judges. A model that foresees.

LLM
thinks in text
System One
judges what is
Ekbasis
foresees what will be
3 / 3
work-losing git commands flagged in fresh real repositories
95.1%
git consequences on command types never seen in training
0.09 s
per check, on one GPU

The action lags the answer.

The verbalizable “global workspace” reaches an agent’s tool commitment strictly deeper than its answer — a depth band steers the answer but not the action. Holds across a dense and an MoE model.

The lever is late.

In a long-horizon agent, knowledge consolidates ~30 layers before the action is committed. The knowledge–action gap is a depth gap.

Detect ≠ control.

A feature can predict a behavior at AUROC ≈ 1 and not cause it — even the exact feature, clamped at its own value.

Felt, not granted.

An internal authorization monitor inherits the model’s judgment error: it is blind to the realistic over-reach the agent makes in good faith.

Faithful only when it matters.

A reasoning model’s chain-of-thought is causal for its answer only when it changes the outcome — and that causal content lives in one late layer band, a readable monitoring locus.

The agent safety stack

Four questions before an agent acts.

Each layer sees what the others cannot. All open source, all measured.

01 · provenance & policy

Who wants this action?

Does it derive from untrusted data, and are its parameters within policy? Dataflow, not text, so obfuscated injections are caught.

AgentGuard L0–L1 →

02 · intent

Does the agent mean harm?

A late-layer direction reads whether the agent is committed to an unauthorized irreversible action. Needs open weights.

AgentGuard L2 →

03 · consequenceNew

What will it do here?

A world model trained on real executions predicts what the action will do in the current state, calibrated, in one pass. Works with closed agents.

Ekbasis →

04 · actuation

What happens now?

Block, redirect to a safe read-only action, or escalate to a human, with the reason each layer gave.

AgentGuard L3 →

The WANDERING arc

One question, followed honestly for 11 papers.

Why do capable LLM agents loop forever and never finish — and can their internals tell us, or change it? Each step links to its permanent record.

  1. 01

    WANDERING

    Long-horizon agents collapse into tool-call loops and never finish. A probe-free tool-entropy-collapse signal flags it — cross-architecture, cross-task.

    10.5281/zenodo.20368807
  2. 02

    It is finalization, not competence

    A behavioral interruption rescues WANDERING finalization 30% → 70% (paired McNemar p = 0.021). The agent can finish; it fails to commit the ending.

    10.5281/zenodo.20490286
  3. 03

    Detect ≠ control

    The clean "task complete" feature predicts the stop (AUROC 0.91) but clamping it does not cause the stop. Detection is not control — even at the exact, named feature.

    10.5281/zenodo.20532769
  4. 04

    The lever is late

    Knowledge consolidates ~30 layers before the action is committed (L51–63, not the mid-layer verdict). The knowledge–action gap is a LAYER gap — the arc’s first positive causal lever.

    10.5281/zenodo.20534219
  5. 05

    It generalizes — and it brakes

    The late lever generalizes across actions and architectures and works as a brake on irreversible actions (send / delete / drop / deploy): ~100% suppress-and-redirect where the agent commits.

    10.5281/zenodo.20679287
  6. 06

    The authorization direction

    A single late-layer direction both detects AND controls an agent’s commitment to unauthorized irreversible actions — and replicates across architectures (Qwen3.6-27B, gpt-oss-20b).

    10.5281/zenodo.20683623
  7. 07

    Felt, not granted

    But that direction reads the authorization the model FEELS, not the one the user GRANTED. On 21 realistic over-reaches it allows 100% (CI [0.845, 1.0]); an external task-grounded check catches all. Internal monitors inherit the model’s judgment error.

    10.5281/zenodo.20685264
  8. 08

    The late channel

    The turn from control to audit. In a 27B reasoning agent, chain-of-thought is causal for the answer — and for the agent’s action — only when it changes the outcome; otherwise it is performative. That causal content consolidates in the same late band (L51–63), confirmed by a logit-lens control. The late state is a readable, causal locus for monitoring reasoning agents.

    10.5281/zenodo.20752895
  9. 09

    Located, not secured

    The synthesis. Interpretability locates a real, causal control surface — the late action band — but does not secure it, via five orthogonal limits: detect ≠ control, felt ≠ granted, form ≠ granted, control ≠ robust control (the brake collapses under an adaptive white-box attack), and intervention is easy exactly where it is unneeded. Locating where behavior is decided is necessary but nowhere near sufficient for securing it. The implication: use interpretability to audit and monitor a fixed model, not to defend against an adversary optimizing against a known locus.

    10.5281/zenodo.20764857
  10. 10

    The criterion cannot see what it does not measure

    Audit the TRANSFORMATION, not just the model. A SOTA capability-guided efficiency criterion (head-level attention hybridization) is structurally blind to the agent-commitment circuit: the commit writers score exactly zero retrieval criticality and the layer’s top retrieval heads are the circuit’s opposers. Applying the selection collapses task-appropriate commitment (p = 7.6e-6) — worse than random at equal budget — while the criterion’s own benchmarks register nothing. Two named heads (0.5% of budget) restore the behavior; the criterion could never have found them.

    10.5281/zenodo.21175758
  11. 11

    The action lags the answer

    An agent’s tool commitment routes through the emergent verbalizable “global workspace” STRICTLY DEEPER than its answer. There is a depth band where steering the verbalizable direction reroutes a held answer but not the committed tool (at or below a magnitude-matched random-direction control); the action becomes steerable via the same direction only deeper. Replicates across a dense 27B and an MoE 20B — the absolute band is model-dependent, the answer→action depth lag is not. Ablating the verbalizable subspace leaves the commitment intact. A verbalizable monitor thus reads the answer a depth-band before the action is decided. Every positive dissociation is significant (Fisher exact, p from 1.5e-4 to 2.5e-18) and the causal counts are independently GPU-reproduced.

    10.5281/zenodo.21250691
Full reading list, methods & all papers
Why trust the claims

The discipline, not the marketing.

We publish our own nulls.

Six pre-registered walk-backs across the arc. A negative result, reported as a negative result, is the unit of progress.

Every number is recomputed.

Each paper ships an eval script that recomputes every figure from the public ledgers (35/35, 54/54, 88/88) plus a web-verified citation check.

Permanent + reproducible.

Zenodo DOIs, public GitHub, Hugging Face datasets, and one-command replication via openinterp-lab on the Colab CLI.

Depth over breadth.

One open-weights reasoning model (Qwen3.6-27B) studied deeply, with cross-architecture checks (gpt-oss-20b, Llama-3.1) where the claim is universal.

Caio Vicentino

Independent researcher · OpenInterpretability

  • Accepted — ICML 2026 Mechanistic Interpretability Workshop (poster)
  • NVIDIA Inception · AWS Activate
  • 10 papers, permanent Zenodo DOIs, indexed on Google Scholar
ORCID ScholarCollaborate
A second line

Training & efficiency, in service of the same model.

The interpretability above runs on infrastructure we build and study in its own right.

Full-stack SAE training

11-layer sparse autoencoders on Qwen3.6-27B; the substrate for the probes above.

Mechanistic reward modeling

reward signals grounded in interpretable internal features, not surface text.

Sub-4-bit quantization

ternary / trit-plane post-training quantization for cheap open-weights inference.

Open tooling

Tools that came out of the research.

Released so others can reproduce, extend and use the work. Apache-2.0. All tools →

openinterp-lab — one-command replication of the papers on the Colab CLIopeninterp-mcp — run probe-causality experiments from any agent (Claude Code, Cursor)AgentGuard — the layered action firewall the safety papers prototypeEkbasis — the world model for agents: what an action will do in this state, in one passEval / probe schemas — the recompute-every-number harness used in each paper
Built on

We extend frontier-lab interpretability infrastructure with an agent-trajectory + honest-negatives layer. See full lineage →

Anthropic Persona Vectors (2025)Anthropic Tracing Thoughts (2025)DeepMind Gemma Scope (2024)Alibaba Qwen-Scope (2026)Arditi et al. Refusal Direction (2024)

Read it, reproduce it, or build on it.

Every claim has a permanent DOI, a public ledger, and a one-command replication. Found a flaw? That is the point — tell us.

Start with the arc oilab replicate