NeFut Logo NeFut
中 Admin Login

[CS.AI] Defense-in-Depth for LLMs: Evaluating Memory Gates Against Activation-Induced and Memory-Induced Sycophancy

Published at: 2026-10-07 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

We propose a $2 \times 2$ defense‑in‑depth framework that separates internal activation steering from external memory handling. Sycophancy steering directions are extracted from 100 paired prompts, and four open‑weight models are evaluated on MemSyco‑Bench across 10 steering coefficients and five memory‑defense configurations (1,550 total answers, with defense conditions judged on a fixed 250‑item subsample by three LLM judges). Three configurations are novel: rewriting every memory, a Router Gate that decides to keep, rewrite, or drop each memory, and dropping all memory; the remaining two are MemSyco baselines.

Results show that selective Router Gate filtering preserves MemSyco’s average accuracy far better than complete memory removal, and this advantage persists when models are steered toward sycophancy. For Llama 3.1 8B, applying Router Gate with mild inverse steering ($\alpha = -1.5$) reduces judge‑averaged sycophancy from 35.80% to 31.32% while average accuracy drops only from 43.99% to 43.31%. The reduction trend is consistent across all three judges but does not reach statistical significance (paired $p = 0.08\sim0.63$ on 149 items).

Overall, external memory filtering proves to be the most robust component; current data do not demonstrate that inverse activation steering adds further benefit.

Review: The study offers a systematic defense strategy, with the Router Gate showing promise in balancing accuracy and mitigating sycophancy. Future work could explore stronger internal steering mechanisms.

Original Source: https://arxiv.org/abs/2610.07403

[h] Back to Home