NeFut Logo NeFut
Admin Login

[CS.AI] The Off-Support Barrier: Limitations of Semantic Safety Constraints in Learning Problems

Published at: 2026-08-13 22:00 Last updated: 2026-08-14 00:05
#AI #Machine Learning #LLM #Artificial Intelligence

We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. Formally, if $q$ is the data distribution and $p(cdotmid w)$ the model, the safety predicate $B$ is not measurable with respect to $sigma( ext{model}, q)$, whereas the real log-canonical threshold (RLCT) of singular learning theory (SLT) is. From this non-invariance we derive, as corollaries rather than independent observations:

Original Source: https://arxiv.org/abs/2608.11243

[h] Back to Home