Probe trained / strongest on GSM8K (AUROC 0.908) drops to 0.500 on MATH — difference −0.408.
Cross-task generalization is not free. Probes labeled “hallucination” often transfer poorly between QA, multiple-choice misconception, and knowledge-recall settings.
Training note: Linear LR probe at L55/mid_think of Qwen3.6-27B reasoning-mode generation. v0.2 trains multi-bench (GSM8K + StrategyQA + MATH combined, 455 samples, 45.8% halu rate). Within-bench AUROC 0.908 on GSM8K held-out — improvement vs v0.1 (0.888). But cross-domain transfer FAILS: 0.612 on StrategyQA (commonsense), 0.500 (chance) on MATH (advanced). AUROC degrades with task difficulty. Position-of-faithfulness in deep residual stream is more strongly domain-bound than multi-bench training compensates for. Multi-bench training thesis (which worked for FabricationGuard cross-task at 0.882) does not transfer to reasoning-faithfulness probes. Honest negative-ish result registered as canonical case study of ProbeBench anti-Goodhart norms — both v0.1 and v0.2 numbers reported without spin.