Abstract
Why do ChatGPT-like AIs, despite major architectural and training differences, unexpectedly tip to undesirable content (e.g. harmful, misleading, repetitive) even under deterministic greedy decoding?
We show that a broad class of such tippings is caused by the many-body interactions between tokens (spins) as they cross the finite-layer system. Tipping emerges as a dynamical first passage process between competing output basins. Attention disorder controls the transport toward, away from, or along the basins' boundary. A few-basin reduction yields a closed finite-layer threshold, whose coarse-grained predictions show good agreement across ChatGPT-like families.
These results suggest that a broad class of AI failures represents 'foreseeable engineering risk' rather than inherently unpredictable behavior, with important implications for legal and societal assessments of AI harm.
Blogger's Review: This study delves into the instability of content generation in ChatGPT-like AIs, revealing the core role of many-body interactions in the generation process. Understanding these mechanisms not only aids in improving AI model design but also provides a crucial perspective for assessing the potential risks posed by AI technologies. By viewing these issues as foreseeable engineering risks, society can more effectively address the challenges brought by AI.