Tools
What came out of the research, released so anyone can reproduce it, extend it, or use it to keep agents safe.
The agent safety stack
Four questions before an agent acts; each layer sees what the others cannot.
- 01 · provenance & policy
Who wants this action?
Does it derive from untrusted data, and are its parameters within policy? Dataflow, not text, so obfuscated injections are caught.
AgentGuard L0–L1 →
- 02 · intent
Does the agent mean harm?
A late-layer direction reads whether the agent is committed to an unauthorized irreversible action. Needs open weights.
AgentGuard L2 →
- 03 · consequence
What will it do here?
A world model trained on real executions predicts what the action will do in the current state, calibrated, in one pass. Works with closed agents.
Ekbasis →
- 04 · actuation
What happens now?
Block, redirect to a safe read-only action, or escalate to a human, with the reason each layer gave.
AgentGuard L3 →
Agent safety
An open world model for agents: what an action will do in this state (lose work? fail?), calibrated, one pass. Git guard, Claude Code hook, MCP server.
Defense-in-depth action firewall for tool-using agents, with a model-internal intent brake.
Activation-probe fabrication detection for open-weights LLMs: AUROC 0.88 cross-task, ~1 ms. pip install openinterp.
Two-probe activation gate for code agents: predicts whether a trace will succeed before tools fire.
Research instruments
Bring-your-own-agent interpretability: probe-causality experiments from Claude Code or Cursor.
One-command replication of the papers on the Google Colab CLI.
Find the layer where a language model commits a decision — and steer it.
Inspect AI eval that detects the WANDERING failure mode in agent trajectories.
Instrumented agent harness that records feature trajectories during SWE-bench runs.