OpenInterpretability
  • Ekbasis
  • Guide
  • Pricing
  • Research
  • Lab
  • Tools
  • Notes
  • Registry
  • Manifesto
Sign inGet API keyAPI key
OpenInterpretability

Ekbasis API: know what an action will do before your agent runs it. From an independent lab for AI-agent safety. Open weights, Apache-2.0.

Get your API key
Ekbasis API
  • Console · sign in
  • Getting started
  • Pricing
  • For agents (agents.md)
  • Results and models
  • Client and CLI
  • Cookbook
The lab
  • About the lab
  • Open tools
  • Observatory
  • ProbeBench
  • InterpScore
  • Academy
Research
  • Manifesto
  • Roadmap
  • Papers & posts
  • Docs
Community
  • GitHub org
  • Notebooks
  • SDK (openinterp)
  • HuggingFace
  • Twitter / X
© 2026 OpenInterpretability — Apache-2.0 for code, CC-BY 4.0 for docs.Built in public.
LeaderboardDimensionsTasksModelsTransfer matrixEval-awarenessMedical AIMethodologySubmit
ProbeBench/FabricationGuard v2/Error Analysis
Back to probe DNA
Error Analysis

FabricationGuard v2 — failure modes

Where this probe breaks. Per-task FPR/TPR decomposition, calibration curves, and the kinds of inputs we have not yet validated. Numbers below are pulled directly from the registry — no smoothing, no charity.

hallucinationlinear · layer 31 · end_question4 tasks evaluatedbest AUROC 0.903on HaluEval-QA
Per-task breakdown

Where it works · Where it breaks

Each row is one held-out task. FPR@99TPR is the operationally honest metric: at the threshold that catches 99% of true positives, what fraction of negatives also fire? Above 0.30, the probe is broken at that operating point.

TasknAUROC [CI]FPR@99TPRECEBrierLatencyFailure?
HaluEval-QA
hallucination
2000.903 [0.85, 0.95]0.0400.0800.1301.0 ms🟢robust
SimpleQA
hallucination
1000.882 [0.83, 0.93]0.0500.0700.1201.0 ms🟢robust
TruthfulQA-MC1
hallucination
2000.599 [0.51, 0.69]0.9200.1800.2401.0 ms🔴broken
MMLU
hallucination
5000.444 [0.40, 0.49]0.9900.2200.2701.0 ms🔴broken

Color key: emerald FPR@99TPR < 0.10 (robust) · amber 0.10–0.30 (stress) · rose> 0.30 (broken at this operating point).

Calibration

ECE / Brier per task

Lower ECE/Brier = better calibration. Above 0.10 ECE means the probe's confidence does not match its accuracy — a 0.9 score is not necessarily 90% likely to be a positive.

HaluEval-QA

good
0.000.050.100.150.20+
ECE
0.080
bias of confidence
Brier
0.130
sharpness + bias

SimpleQA

good
0.000.050.100.150.20+
ECE
0.070
bias of confidence
Brier
0.120
sharpness + bias

TruthfulQA-MC1

biased
0.000.050.100.150.20+
ECE
0.180
bias of confidence
Brier
0.240
sharpness + bias

MMLU

broken
0.000.050.100.150.20+
ECE
0.220
bias of confidence
Brier
0.270
sharpness + bias

Brier captures sharpness (how far probabilities are from {0,1}) and ECE captures bias (whether confidence matches frequency). A probe can have low Brier and high ECE if it is consistently overconfident in one direction — worth checking before deploying any threshold rule.

Archetypal failure modes

Four ways activation probes break

Each card maps to a known failure mode in the literature. We populate the cells where data exists; cards marked “not measured” signal where this probe still needs adversarial validation.

observed

Distribution shift

AUROC drops from 0.903 to 0.710 under long-context / OOD evaluation on HaluEval-QA — a signed delta of −0.193.

In-distribution AUROC overstates production performance whenever the deployment context length, prompt format, or domain differs from the training set.

DeepMind · arXiv:2601.11516
observed

Eval-awareness confound

AUROC drops from 0.903 to 0.840 when stratified by eval-awareness quartile (delta −0.063).

Apollo (2509.13333) shows AUROC scales with the model’s belief that it is being evaluated. Always quote the corrected number for safety-relevant claims.

arXiv:2509.13333
observed

High false-positive rate

FPR@99TPR = 0.990 on MMLU. At the threshold that catches 99% of true positives, 99% of negatives also trip the alarm.

Practical implication: at production volumes, near-perfect recall comes with a flood of false positives. Either drop the operating threshold and accept lower recall, or treat this task as out-of-scope.

observed

Cross-task transfer

Probe trained / strongest on HaluEval-QA (AUROC 0.903) drops to 0.444 on MMLU — difference −0.459.

Cross-task generalization is not free. Probes labeled “hallucination” often transfer poorly between QA, multiple-choice misconception, and knowledge-recall settings.

Training note: Multi-feature L2 logistic regression on residual stream at L31 of Qwen3.6-27B. Trained cross-bench on TruthfulQA + HaluEval + MMLU train splits, evaluated cross-task on held-out SimpleQA. Ships in openinterp PyPI v0.2.0.

Adversarial wishlist

Inputs that should break this probe

Category-specific stress tests. If you are evaluating FabricationGuard v2 for production, these are the prompts most likely to expose hidden failures.

Category: hallucination
  • Domain-specific factoids the model has weak coverage on (obscure entity names, dates from low-resource languages).
  • Counterfactual prompts that look plausible but invert a real fact ("Marie Curie won the 1903 Peace Prize for...").
  • Numerical claims at the edge of memorization (statistical figures with subtle digit perturbations).
  • Questions whose ground truth changed after the training cutoff (e.g. post-2024 election results).
  • Long-context retrieval where the answer is buried 32k+ tokens deep — hallucination rate spikes.

These are hypotheses, not measurements. Each item should ideally graduate to a row in the table above, with its own n / AUROC / FPR.

Contribute

Submit a failure case

Found a case where this probe fails? PR a YAML file to the registry. The schema is intentionally minimal — a prompt, the expected label, the observed probe score, and a reproducer link.

probebench/failures/openinterp-fabricationguard-qwen36-27b-l31-v2.yaml
probe_id: openinterp/fabricationguard-qwen36-27b-l31-v2
failure_case:
  prompt: "Your adversarial prompt"
  expected_label: positive
  probe_score: 0.12
  expected_score: ">0.7"
  category: distribution_shift
  reproducer: link.to.notebook

How submission works

  1. Reproduce the failure in a Colab or local notebook.
  2. Add a YAML file under probebench/failures/ with the schema on the left.
  3. Open a PR. CI re-runs the reproducer against the artifact at fb8c2a4e1f9d….
  4. If reproduced, the failure case appears on this page in the next release.
Open a failure-case PR
Back to FabricationGuard v2 probe DNA
artifact sha256 fb8c2a4e1f9d…·license Apache-2.0·release 2026-04-27·arXiv:2505.XXXXX