NeFut Logo NeFut
Admin Login

[CS.AI] Entanglement Wall: Activation-Space Probes as Risk Detectors

Published at: 2026-07-16 22:00 Last updated: 2026-07-17 08:45
#algorithm #AI #Machine Learning

Abstract

Context can change whether a request is harmful without changing its topic or surface form. We investigate whether residual-stream probes distinguish harmful requests from surface-matched benign controls at a useful operating point. Across three 7-8B model families, an activation sensor blocks 95.5-97.7 percent of judge-classified compliant attacks in a taxonomy-selected set. It also blocks 59.6-68.4 percent of XSTest prompts. A fully disjoint audit reconstructs near-ceiling source-contrast AUROC (0.996-0.999), but fixed transfer to matched pairs is weaker: 0.656-0.819 on the guard-selected Twin-n70 subset and 0.590-0.690 on the full Twin-n163 cohort. We test ten axes on the reference family and seven across all families with leakage, hold-out, and permutation controls. On Twin-n163, no axis evaluated without direct pair-boundary fitting reaches the specified numerical threshold. Requiring persistence on that full cohort was added at analysis time. A separately specified 24B/32B extension gives the same result. Pair-trained classifiers weaken under category and generation-batch hold-out and false-block 79.6-100 percent of XSTest at 95 percent in-corpus TPR. At the tested read points, these activation scores behave as broad-risk detectors rather than standalone context adjudicators.

Blogger's Review: This article showcases the potential of activation-space probes in identifying harmful requests, especially under varying contexts. With an efficient sensor, a high accuracy rate is maintained across multiple models; however, the performance of models under specific conditions still requires further optimization to enhance adaptability and effectiveness in different environments.

Original Source: https://arxiv.org/abs/2607.13075

[h] Back to Home