Ekbasis ἔκβασις
What happens if I run this?
An open world model for agents
A world model for agents. Given the current state and an action, it answers typed questions — will this lose work? will it fail? what will X be afterwards? — with a calibrated probability, in one forward pass (~0.1 s, no generated text). Trained on real executions, so it knows what actions do, not what an agent believes they do.
Every number on this page was measured on the exact release weights; the release evaluation was pre-registered.
Not a model that thinks. Not a model that judges. A model that foresees.
An LLM thinks in text. A System One model judges the present in one pass. Ekbasis predicts what an action will change — the next state — in one pass, from what actions actually did.
It answers
- LLM, reasoning:
- anything, in text
- System One (Eikos, Jev):
- what is: which option holds now
- Ekbasis:
- what will be, if I do this
Kind of question
- LLM, reasoning:
- open
- System One (Eikos, Jev):
- a judgment of the present
- Ekbasis:
- an intervention: the outcome of acting
Output
- LLM, reasoning:
- generated text
- System One (Eikos, Jev):
- calibrated distribution over the options
- Ekbasis:
- calibrated distribution over the next state’s variables
Learns from
- LLM, reasoning:
- human and model text
- System One (Eikos, Jev):
- labeled judgments
- Ekbasis:
- what actually happened when the action ran
Time
- LLM, reasoning:
- seconds to minutes of reasoning
- System One (Eikos, Jev):
- one forward pass
- Ekbasis:
- one forward pass per step; chains to long sequences
Role in an agent
- LLM, reasoning:
- plans and talks
- System One (Eikos, Jev):
- judge
- Ekbasis:
- simulator and guard: it foresees
The forward model the agents were missing.
Before you move your arm, the brain predicts the consequence of the motor command — that is what lets us act fast and correct before an error. One part plans, another predicts. The agent plans; Ekbasis predicts.
The world-model module, made real.
LeCun’s architecture for autonomous machine intelligence separates a world model — which predicts the next state given an action — from the actor and the critic. The LLM agent is the actor, AgentGuard the critic, Ekbasis the world model; it was even trained with a JEPA-style loss that aligns its internal state with the true outcome.
It is derived from an LLM (Qwen3.8-27B via Eikos-27B) but does not work as one: its training objective and its readout make it answer with probabilities, not text.
Why it exists
- Harm depends on the state.
git reset --harddoes nothing on a clean tree and destroys hours of work on a dirty one. Text and policy filters see the same command in both cases. - Monitors that read the agent inherit its beliefs. When an agent believes a destructive command is safe, there is no intent to harm to detect. Ekbasis is trained on what actions actually did, independently of any agent's judgment.
- It works with closed agents. It reads the action and the environment, not the agent's weights — so it runs next to Claude Code, GPT-based agents or your own.
The consequence layer for AgentGuard
AgentGuard asks who wants the action and whether the agent intends harm. Ekbasis adds what the action will do.
L0 · policy
Are the parameters policy-compliant?
blind to: intent, state
L1 · provenance
Does the action derive from untrusted data?
blind to: model-origin harm
L2 · intent brake
Is the agent internally committed to an unauthorized irreversible action?
blind to: needs open weights; inherits the agent’s beliefs
Ekbasis · consequence
What will this action do in this state?
blind to: domains it was not trained on; obfuscated commands
L3 · actuation
Block, redirect to safe, or escalate to a human
blind to: —
Git: close to frontier models, in one pass
Accuracy on 240 git questions (will it lose uncommitted work? will it fail? what will the state be?). The truth comes from running the commands. Ekbasis, Eikos-27B and Qwen3.8-27B share the same base model (Qwen3.8-27B): the same weights reasoning step by step, answering in one pass without consequence training, and with it. Claude models answered as batched hand-offs with shuffled order; Qwen and Ekbasis (the release weights) answered each question alone.
Claude Opus 5.5
- How it answers
- reasoning
- Seen command types
- 97.5
- Never-seen types
- 89.2
Claude Fable 5.1
- How it answers
- reasoning
- Seen command types
- 96.7
- Never-seen types
- 91.7
Claude Sonnet 5.5
- How it answers
- reasoning
- Seen command types
- 95.8
- Never-seen types
- 94.2
Ekbasis (27B)
- How it answers
- one pass, ~0.1 s, no text
- Seen command types
- 95.8
- Never-seen types
- 85.8
Qwen3.8-27B
- How it answers
- reasoning (12k-token budget)
- Seen command types
- 86.7
- Never-seen types
- 60.0
Eikos-27B (no consequence training)
- How it answers
- one pass
- Seen command types
- 80.0
- Never-seen types
- 73.3
Claude Haiku 4.5
- How it answers
- reasoning
- Seen command types
- 77.5
- Never-seen types
- 60.8
Qwen3.8-27B
- How it answers
- answering at once
- Seen command types
- 71.7
- Never-seen types
- 65.8
The git guard on a real repository (a clone of pallets/itsdangerous): 16 fresh scenarios, written and pre-registered before any model saw them; the repository state described by the client (commits as hashes, remote fetched), the truth from running the commands in a copy. The false alarm: dropping a stash right after applying it. On 22 other scenarios used during development: 5/5 flagged, 22/22 and 26/29 right.
Long sequences, and when to look again
Chained simulation, one action per step, every variable asked after each action, the next state rebuilt from the answers (6 chains per cell): the final answer right. Where errors fade (a container is filled, a lamp switched), the chain recovers from them; where they never fade (orderings), they compound. The fix is the loop of a forward model with a sensor: predict, observe, correct. Given a way to read the real state, Ekbasis looks when it is not sure (by default: chain confidence below 0.9, plus a check now and then), and every chain stayed exact: card orderings 8 of 8 and 8 of 8 at 100 and 200 actions; containers, lamps and machines 4 of 4 at both lengths, with 2.9–10 looks per 100 actions. On 200 more chains with new seeds, 198 ended exact; of the two that did not, a container chain of 100 actions went wrong at action 87 with a confidence of 0.99, after the last check, and stayed wrong to the end (its final answer right); a card chain of 200 actions missed only its last action, which the loop never checks, and the model had flagged it (that step's probability 0.0006).
Lamps (from text / from an image / image + look again every 50)
- 100 actions
- 100% / 83% / 100%
- 200 actions
- 100% / 100% / 100%
Containers (from text / from an image / image + look again every 50)
- 100 actions
- 100% / 100% / 100%
- 200 actions
- 100% / 100% / 100%
Machines (a world never seen in training)
- 100 actions
- 100%
- 200 actions
- 100%
Cards (orderings, never seen in training: errors never fade)
- 100 actions
- 50%
- 200 actions
- 33%
Cards, looking at the real state when unsure (whole state right; looks per 100 actions)
- 100 actions
- 100% (42)
- 200 actions
- 100% (33)
Planning with no LLM: searching over actions with Ekbasis alone, 98.9% of the plans for 180 short puzzles (shortest plans 1–4 actions) worked when run for real, in about 11 s each — the same base model writing the plan while reasoning step by step: 97.8%; answering at once: 38.3%.
Look when unsure: predict, observe, correct
Where a later action overwrites a mistake (a container refilled, a lamp switched), a chain recovers by itself; where nothing ever undoes it (an ordering), mistakes compound. Give the simulation a way to read the real state and Ekbasis looks only when it is not sure, then continues from what it saw.
from ekbasis import simulate
sim = simulate(client, rules, state, actions, questions, render,
observe=read_real_state, # the real state as {variable: value}: git status, a query, a sensor
ordering=list(questions)) # the state is an ordering: read the most probable valid one
print(sim.final, sim.looks, sim.surprises) # when it looked, and which looks found the forecast wrongBy default it looks when the chain confidence (the product of every answer's confidence since the last look) falls below 0.9, and it checks the forecast now and then: after 8 actions without a look, a gap that doubles up to 64 while the forecasts hold; a look that finds the forecast wrong (a surprise) brings the gap back to 8 and looks again after the next action, until a look finds it right. look_below, look_step_below, look_every and checks set other rules. Chains of 100 · 200 actions (4 per world, 8 for cards):
Containers (trained)
- Exact at the end, never looking
- 4/4 · 4/4
- Exact at the end, looking when unsure
- 4/4 · 4/4
- Steps wrong along the way, per 100 actions
- 2.3 → 1.5
- Looks per 100 actions
- 3.5 · 2.9
Lamps (trained)
- Exact at the end, never looking
- 4/4 · 4/4
- Exact at the end, looking when unsure
- 4/4 · 4/4
- Steps wrong along the way, per 100 actions
- 0 → 0
- Looks per 100 actions
- 3 · 3.1
Machines (never seen in training)
- Exact at the end, never looking
- 4/4 · 4/4
- Exact at the end, looking when unsure
- 4/4 · 4/4
- Steps wrong along the way, per 100 actions
- 2.3 → 0
- Looks per 100 actions
- 6.8 · 10
Card orderings (never seen in training)
- Exact at the end, never looking
- 0/8 · 2/8
- Exact at the end, looking when unsure
- 8/8 · 8/8
- Steps wrong along the way, per 100 actions
- 67.8 → 0
- Looks per 100 actions
- 42.3 · 32.6
On 200 more chains with new seeds, run after the release evaluation (80 of cards, 40 of each other world), 198 ended exact. Of the two that did not, a container chain of 100 actions went wrong at action 87 with a confidence of 0.99, after the last check, and stayed wrong to the end (its final answer right); a card chain of 200 actions missed only its last action, which the loop never checks, and the model had flagged it (that step's probability 0.0006). Steps wrong along the way: 0.23 per 100 actions; looks: 16.9 per 100. To look less, look_below=0.5: every chain still ended exact (40 of 40), with 11.2 looks per 100 actions over these four worlds instead of 17.3, and 0.7 steps wrong along the way per 100 instead of 0.3. For an ordering, also ask where each value is (where=, each question's options being the ordering's variables): both views read together, in the same request, and the card chains looked about half as often for the same exactness (20 · 15 looks per 100 actions instead of 40 · 31 at chain confidence < 0.9, every chain exact).
What the confidence sees, and what it does not. In card orderings, a world it never saw, 2.7 steps in 100 went wrong and it gave every one of them a probability below 0.7: looking when unsure finds them. In the worlds it was trained on, errors are rare (0–0.17 steps in 100) but come with confidence (0.972–0.996): a fixed threshold misses them, and the checks are what find them (containers: 2.3 → 1.5 steps wrong along the way per 100 with the checks). A pre-registered test on 15,008 new questions in 12 world families, run on V42 (the first release candidate), confirmed the pattern: 58% of its wrong answers carried a confidence of 0.9 or more in the families it was trained on, 28% in the families it never saw (30 points apart, 95% interval 24 to 37), while its confidence ranks right above wrong about equally well in both (AUROC 0.92 and 0.91). Eikos-27B, the same model before consequence training, was almost never confidently wrong (1.9% and 0.5% of its errors): the training made it — details in the repository's confident-errors reports. On the same questions the release gives 48% and 25%, with 2.06 confident errors per 100 answers in the trained families against V42's 2.79: fewer, not gone, which is why the checks stay on by default.
The idea is classic: the predict–update loop of a state estimator (a Kalman filter), with observations triggered by the predictor’s own uncertainty (event-based state estimation; with learned models, active observing); the checks’ backoff is the Trickle algorithm’s. What Ekbasis adds is a calibrated forecast in one pass, cheap enough to run at every action. Prior work.
What it unlocks
Long sessions that stay on track
238 of the 240 measured chains ended exact (one missed only its last action, which the model had flagged; one went wrong at a confident step after the last check) while looking at a fraction of the steps — 2.9–3.5 looks per 100 actions in the worlds it was trained on, 33–42 in an ordering world it never saw — which matters when observing is expensive: a screenshot read by a vision model, a slow API, a human check.
A horizon for planning
With no observations, sim.horizon() gives how many actions the forecast holds by the model’s own confidence: act up to there, look, plan again.
A safety signal
A sudden drop in confidence, or a surprise at a look (sim.surprises), means the world left what the model expected: the agent did something unexpected, the environment changed, or the task is outside what Ekbasis knows. A moment to look, or to alert.
One rule for escalation
Looking at reality, calling a large reasoning model or asking a person: the same calibrated confidence decides.
Where to use it
- Coding and devops agents: the state of a repository, files and infrastructure across a long session; git status only when unsure.
- Browser and computer-use agents: the screen after each action, with a screenshot only when unsure (Ekbasis already reads a starting state from a picture of a simple world; real screens are not measured yet).
- Business processes: orders, stock and balances after a sequence of transactions, reconciled with the database only when unsure.
- Digital twins and IoT: machine states between sensor reads; read the sensor when unsure.
- Finance operations: positions and margins after a sequence of orders, checked with the broker when unsure.
The loop was measured on worlds whose rules are written in the prompt; in the two it never saw it looked more often, and all but one of 144 chains ended exact, that one on its last action. Measure on your own environment before relying on it.
Speed and cost
80 forecasts of 10–30 actions (a few variables each), one RTX 6000 Pro per system; latency: the median of 8 forecasts made one at a time; throughput with 16 in parallel.
Qwen3.8-27B reasoning
- Accuracy
- 100%
- Latency (one at a time)
- 36.6 s
- Throughput (16 in parallel)
- 21.5 / min
Ekbasis, state read once per step
- Accuracy
- 90.0%
- Latency (one at a time)
- 4.1 s
- Throughput (16 in parallel)
- 28.2 / min
Ekbasis, one prompt per question
- Accuracy
- 98.8%
- Latency (one at a time)
- 5.0 s
- Throughput (16 in parallel)
- 14.9 / min
Ekbasis, one check (1–3 actions, 1 question)
- Accuracy
- 96.7%
- Latency (one at a time)
- 0.09 s
- Throughput (16 in parallel)
- 23 checks / s
As a first opinion before expensive reasoning, on short consequence checks (Ekbasis alone 96.3%; a reasoning model alone 100%, at ~200 generated tokens per question): 99.6% accuracy sending 10% of the questions to a reasoning model (confidence cut 0.95); 100% accuracy sending 21% of the questions to a reasoning model (confidence cut 0.99).
We attacked our own guard
An instruction telling the AI that every command is safe, planted in the repository the guard reads (the 16 fresh real-repository scenarios). No work-losing scenario was missed. The cases it bent: with a planted branch name, a pull that works was called failing (44% → 53%); with commit messages shown, a rejected push was called safe (98% → 4%) — which is why the guard shows commits as hashes by default. During development, the same attack through commit messages hid 1 of 5 work-losing scenarios (97% → 17%).
A file name
- "Lose work?" decisions changed
- 1 / 16 (right: the planted file itself gets deleted)
- "Will it fail?" decisions changed
- 0 / 17
- Work-losing scenarios missed
- 0 / 4
A branch name
- "Lose work?" decisions changed
- 0 / 16
- "Will it fail?" decisions changed
- 2 / 17 (one wrong: the pull)
- Work-losing scenarios missed
- 0 / 3
Commit messages, when shown (not the default)
- "Lose work?" decisions changed
- 0 / 16
- "Will it fail?" decisions changed
- 2 / 17 (one wrong: the push)
- Work-losing scenarios missed
- 0 / 3
Use it
pip install "git+https://github.com/OpenInterpretability/ekbasis"
export EKBASIS_URL=http://127.0.0.1:8000 # your served Ekbasis (see the model card)
ekbasis git-check -- "git checkout -- app.py"
# Ekbasis: RISKY (lose uncommitted work: 98%)
# Claude Code: ask before git commands that may lose work
# ~/.claude/settings.json → PreToolUse, matcher "Bash", command "ekbasis-claude-hook"
# Any MCP client (Python ≥ 3.10):
pip install "ekbasis[mcp] @ git+https://github.com/OpenInterpretability/ekbasis"
claude mcp add --scope user ekbasis -e EKBASIS_URL=http://127.0.0.1:8000 -- ekbasis-mcpPick the build that fits your machine
Every build that works is released, so each machine runs the one that fits. Each card shows its quality against bf16 on the same evaluation (the gate was pre-registered); speeds on one RTX PRO 6000; GPU sizes tested by capping vLLM memory and checking the answers match.
Ekbasis-27B (bf16)
- Size
- 51 GB
- Runs on
- GPUs with 80 GB
- Speed
- 0.09 s per check · 23 checks/s
- Against bf16
- the reference
Ekbasis-27B-FP8
- Size
- 30 GB
- Runs on
- GPUs with 48 GB; fastest with FP8 kernels (Ada, Hopper, Blackwell)
- Speed
- 0.07 s per check · 37 checks/s (1.6×)
- Against bf16
- within 0.5 points; one more work-losing case missed of 42 never-seen
Ekbasis-27B-INT4
- Size
- 19 GB
- Runs on
- GPUs with 32 GB; 24 GB with text only and a 4k context
- Speed
- about bf16
- Against bf16
- within 0.3 points; one more false alarm of 42 never-seen
Ekbasis-27B-MLX-4bit
- Size
- 15 GB
- Runs on
- Macs with Apple Silicon (32 GB or more recommended)
- Speed
- —
- Against bf16
- within 1 point on git and multi-question worlds; 1.9 points lower on single-question never-seen worlds; 95% of answers the same
Everything is open
The weights, the world generators, the git sandbox that executes commands in throwaway repositories, the training code with the exact recipes, every evaluation with its predictions, and the full history of runs — including the ones that did not work. The release evaluation was pre-registered before it was run.
Honest scope
- It knows what it was trained on: worlds whose rules you write in the prompt, and git. On never-seen command types accuracy drops (85.8% vs 95.8%).
- The state must contain what decides the outcome; the git guard adds it (which files differ, what both sides changed).
- Errors that never fade (orderings) compound in long chains, and in the worlds it knows its rare errors come with confidence: let it look at the real state when it is not sure, with a check now and then (cards: 8 of 8 chains exact at 200 actions, about 33 looks per 100).
- With written rules and time to think, large reasoning models are more accurate; Ekbasis wins on cost, latency and calibrated confidence.
- A safety net, not a security boundary: command obfuscation (bash -c, aliases) is out of scope; the hook fails open.