The current debate on machine consciousness rests on a hidden premise: that the entity labeled "AI" already qualifies as a potential bearer of consciousness. This paper first separates phenomenal consciousness, introspective reports, and human projective introspection, arguing that generative systems can produce first‑person linguistic traces of human interiority without actually possessing a phenomenal bearer.
From this observation the AI Consciousness Fallacy is defined: mistaking a system's output for evidence of consciousness. The authors then introduce Causal Liability Theory (CLT). CLT‑I proposes liability closure as a criterion for individuating a candidate bearer: a physically continuous process becomes the non‑delegable inheritor of constraints generated by its own endogenous discriminations.
CLT‑II extends the claim, positing that liability closure is both necessary and sufficient for minimal phenomenal subjecthood. To test these claims, an open‑weight causal audit is performed across several model families.
Key experimental observations include:
- Forced discriminations produce persistent downstream divergence;
- Activation patching reveals strong causal mediation;
- Live and copied adaptive states are behaviorally identical when randomness is matched;
- Detached reconstruction preserves computational state while, by protocol, breaking continuity and non‑delegable inheritance.
The results demonstrate that CLT‑I distinctions are experimentally tractable and can separate causal bearer structure from first‑person performance. Consequently, the framework isolates consciousness attribution, causal bearer individuation, and the metaphysical question of consciousness constitution as distinct layers.
Review: The paper offers a concrete experimental methodology that cautions against equating linguistic output with genuine subjective experience, and it advances a formal approach to causal responsibility in AI systems.