Back homeOPEN SOURCE · APACHE-2.0

Tools

What came out of the research, released so anyone can reproduce it, extend it, or use it to keep agents safe.

The agent safety stack

Four questions before an agent acts; each layer sees what the others cannot.

  1. 01 · provenance & policy

    Who wants this action?

    Does it derive from untrusted data, and are its parameters within policy? Dataflow, not text, so obfuscated injections are caught.

    AgentGuard L0–L1 →

  2. 02 · intent

    Does the agent mean harm?

    A late-layer direction reads whether the agent is committed to an unauthorized irreversible action. Needs open weights.

    AgentGuard L2 →

  3. 03 · consequence

    What will it do here?

    A world model trained on real executions predicts what the action will do in the current state, calibrated, in one pass. Works with closed agents.

    Ekbasis →

  4. 04 · actuation

    What happens now?

    Block, redirect to a safe read-only action, or escalate to a human, with the reason each layer gave.

    AgentGuard L3 →

Agent safety

Research instruments

Benchmarks & standards

Training