NeFut Logo NeFut
中 Admin Login

[CS.AI] ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation

Published at: 2026-09-26 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

ENDOPROMPT is a white‑box approach that learns utility‑degrading prefixes from unlabeled instructions. The generator takes the request text as input, while clean victim continuations serve as pseudo‑references. A local search first identifies prefixes that reduce the likelihood of the continuation; these prefixes are then used for preference fitting on comparisons within the same instruction, followed by reward refinement that distills the signal into the generator. At deployment the generator emits one prefix per request without any further victim‑side search.

Across four instruction‑tuned models and the full splits of seven benign benchmarks, ENDOPROMPT achieves an average utility drop of 26.8 percentage points, with 27 out of 28 cells showing negative change. Failure analysis points to output expansion and prefix reuse as main issues, and control experiments indicate that matching the request does not provide a degradation advantage. The results demonstrate that supervision derived from the victim can expose utility weaknesses without benchmark feedback or predefined failure responses. The code will be released upon acceptance.

Review

Original Source: https://arxiv.org/abs/2609.29948

[h] Back to Home