An independent lab studying how LLM agents fail on long-horizon tasks — and building the open tools that catch those failures before they do harm.
We publish positives and nulls with the same rigor: pre-registered, every number recomputed from public data, permanent Zenodo DOIs. The WANDERING arc explains how agents fail; AgentGuard and Ekbasis stop the failures that matter.
An open world model for agents
Before an action runs, it predicts what the action will do — will this lose work? will it fail? — with a calibrated probability, in one forward pass. Trained on what actions actually did.
Meet EkbasisNot a model that thinks. Not a model that judges. A model that foresees.
The action lags the answer.
The verbalizable “global workspace” reaches an agent’s tool commitment strictly deeper than its answer — a depth band steers the answer but not the action. Holds across a dense and an MoE model.
The lever is late.
In a long-horizon agent, knowledge consolidates ~30 layers before the action is committed. The knowledge–action gap is a depth gap.
Detect ≠ control.
A feature can predict a behavior at AUROC ≈ 1 and not cause it — even the exact feature, clamped at its own value.
Felt, not granted.
An internal authorization monitor inherits the model’s judgment error: it is blind to the realistic over-reach the agent makes in good faith.
Faithful only when it matters.
A reasoning model’s chain-of-thought is causal for its answer only when it changes the outcome — and that causal content lives in one late layer band, a readable monitoring locus.
Each layer sees what the others cannot. All open source, all measured.
Who wants this action?
Does it derive from untrusted data, and are its parameters within policy? Dataflow, not text, so obfuscated injections are caught.
AgentGuard L0–L1 →
Does the agent mean harm?
A late-layer direction reads whether the agent is committed to an unauthorized irreversible action. Needs open weights.
AgentGuard L2 →
What will it do here?
A world model trained on real executions predicts what the action will do in the current state, calibrated, in one pass. Works with closed agents.
Ekbasis →
What happens now?
Block, redirect to a safe read-only action, or escalate to a human, with the reason each layer gave.
AgentGuard L3 →
Why do capable LLM agents loop forever and never finish — and can their internals tell us, or change it? Each step links to its permanent record.
Long-horizon agents collapse into tool-call loops and never finish. A probe-free tool-entropy-collapse signal flags it — cross-architecture, cross-task.
A behavioral interruption rescues WANDERING finalization 30% → 70% (paired McNemar p = 0.021). The agent can finish; it fails to commit the ending.
The clean "task complete" feature predicts the stop (AUROC 0.91) but clamping it does not cause the stop. Detection is not control — even at the exact, named feature.
Knowledge consolidates ~30 layers before the action is committed (L51–63, not the mid-layer verdict). The knowledge–action gap is a LAYER gap — the arc’s first positive causal lever.
The late lever generalizes across actions and architectures and works as a brake on irreversible actions (send / delete / drop / deploy): ~100% suppress-and-redirect where the agent commits.
A single late-layer direction both detects AND controls an agent’s commitment to unauthorized irreversible actions — and replicates across architectures (Qwen3.6-27B, gpt-oss-20b).
But that direction reads the authorization the model FEELS, not the one the user GRANTED. On 21 realistic over-reaches it allows 100% (CI [0.845, 1.0]); an external task-grounded check catches all. Internal monitors inherit the model’s judgment error.
The turn from control to audit. In a 27B reasoning agent, chain-of-thought is causal for the answer — and for the agent’s action — only when it changes the outcome; otherwise it is performative. That causal content consolidates in the same late band (L51–63), confirmed by a logit-lens control. The late state is a readable, causal locus for monitoring reasoning agents.
The synthesis. Interpretability locates a real, causal control surface — the late action band — but does not secure it, via five orthogonal limits: detect ≠ control, felt ≠ granted, form ≠ granted, control ≠ robust control (the brake collapses under an adaptive white-box attack), and intervention is easy exactly where it is unneeded. Locating where behavior is decided is necessary but nowhere near sufficient for securing it. The implication: use interpretability to audit and monitor a fixed model, not to defend against an adversary optimizing against a known locus.
Audit the TRANSFORMATION, not just the model. A SOTA capability-guided efficiency criterion (head-level attention hybridization) is structurally blind to the agent-commitment circuit: the commit writers score exactly zero retrieval criticality and the layer’s top retrieval heads are the circuit’s opposers. Applying the selection collapses task-appropriate commitment (p = 7.6e-6) — worse than random at equal budget — while the criterion’s own benchmarks register nothing. Two named heads (0.5% of budget) restore the behavior; the criterion could never have found them.
An agent’s tool commitment routes through the emergent verbalizable “global workspace” STRICTLY DEEPER than its answer. There is a depth band where steering the verbalizable direction reroutes a held answer but not the committed tool (at or below a magnitude-matched random-direction control); the action becomes steerable via the same direction only deeper. Replicates across a dense 27B and an MoE 20B — the absolute band is model-dependent, the answer→action depth lag is not. Ablating the verbalizable subspace leaves the commitment intact. A verbalizable monitor thus reads the answer a depth-band before the action is decided. Every positive dissociation is significant (Fisher exact, p from 1.5e-4 to 2.5e-18) and the causal counts are independently GPU-reproduced.
Six pre-registered walk-backs across the arc. A negative result, reported as a negative result, is the unit of progress.
Each paper ships an eval script that recomputes every figure from the public ledgers (35/35, 54/54, 88/88) plus a web-verified citation check.
Zenodo DOIs, public GitHub, Hugging Face datasets, and one-command replication via openinterp-lab on the Colab CLI.
One open-weights reasoning model (Qwen3.6-27B) studied deeply, with cross-architecture checks (gpt-oss-20b, Llama-3.1) where the claim is universal.
Independent researcher · OpenInterpretability
The interpretability above runs on infrastructure we build and study in its own right.
11-layer sparse autoencoders on Qwen3.6-27B; the substrate for the probes above.
reward signals grounded in interpretable internal features, not surface text.
ternary / trit-plane post-training quantization for cheap open-weights inference.
Released so others can reproduce, extend and use the work. Apache-2.0. All tools →
We extend frontier-lab interpretability infrastructure with an agent-trajectory + honest-negatives layer. See full lineage →
Every claim has a permanent DOI, a public ledger, and a one-command replication. Found a flaw? That is the point — tell us.