NeFut Logo NeFut
中 Admin Login

[CS.AI] Frequency Is Not Sensitivity: Identifying Safety‑Sensitive Experts in Sparse MoE LLMs

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Suppressing a small set of routed experts can weaken the safety behavior of a sparse Mixture‑of‑Experts (MoE) language model without retraining, making the choice of experts a security question. The usual answer relies on activation frequency, which measures usage rather than influence.\ We evaluate an alternative metric—router‑gradient sensitivity, i.e., the sensitivity of the sequence loss to the gate weights that select an expert.\ Across five MoE architectures we rank experts using 500 benign and 500 malicious prompts, then measure refusal on 100 held‑out malicious prompts under two budget constraints: equal expert counts and equal nominal malicious routing traffic (1%‑5%).\ Under each budget, router‑gradient selection reduces refusals more than activation in 24 of 25 conditions and outperforms the mean of ten random trials in all 25.\ The strongest effect appears in OLMoE, where refusals drop from 34/100 to 9/100 (a 73.53% relative reduction) with no degradation in output quality, indicating genuine compliance rather than broken generation.\ Even after matching expert counts in every layer, gradient selection still yields greater refusal reduction in 23 of 25 conditions (two ties).\ An exploratory cross‑model analysis links larger malicious‑versus‑benign concentration gaps to stronger peak gradient effects (ρ = 0.90; exact two‑sided p = 0.083, n = 5).\ Overall, the results support using router‑gradient sensitivity for expert suppression under the tested budgets.\ \ Review: Router‑gradient sensitivity offers a causally grounded safety control that surpasses simple activation frequency, markedly improving refusal rates in high‑risk scenarios without sacrificing generation quality.

Original Source: https://arxiv.org/abs/2610.02910

[h] Back to Home