NeFut Logo NeFut
中 Admin Login

[CS.AI] Chain-of-Thought Monitorability of Looped Language Models

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Chain‑of‑thought (CoT) monitoring is a promising way to catch undesirable model behavior. Looped language models (LoopLMs) repeatedly apply the same Transformer layers, increasing effective computational depth without adding parameters. This work provides the first systematic evaluation of CoT monitorability for LoopLMs. We consider two complementary setups: (1) fixing the model family and varying the loop depth to isolate the effect of extra recurrent computation; (2) comparing LoopLMs with non‑looped models matched on parameter count, layer number, or effective depth to see whether the looped architecture harms monitorability. Experiments cover eight tasks from MonitorBench under both standard and stress‑test conditions. We observe that, on certain Logic/Science/Engineering “Cue Answer” tasks, deeper loops reduce CoT monitorability in stress tests, while other tasks show weaker or qualitatively different patterns. Diagnostic analysis indicates that the drop cannot be fully explained by task difficulty, verification pass rate, or token length, and deeper‑loop models change how they explicitly use or attribute the provided cues. Cross‑model comparison finds no systematic evidence that LoopLMs are less monitorable than size‑ or depth‑matched non‑looped models. In summary, increasing loop depth can impair CoT monitorability on some stress‑test tasks, but the looped Transformer design alone does not necessarily imply lower monitorability.

Review

Original Source: https://arxiv.org/abs/2610.02741

[h] Back to Home