NeFut Logo NeFut
Admin Login

[CS.AI] Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts

Published at: 2026-09-02 22:00 Last updated: 2026-09-03 02:56
#AI #Machine Learning #LLM

Large Language Models (LLMs) can solve complex problems, yet misuse in high‑risk domains may cause severe damage. Model providers therefore restrict assistance for potentially harmful requests. Refusing all cybersecurity queries would harm legitimate users, so providers need a way to block malicious use without denying defenders legitimate help. Existing cybersecurity‑specific datasets evaluate this mechanism but ignore the conversational context of a request. We introduce 3R‑Bench (Refusal, Repetition, and Revision), a benchmark of 150 real‑world cybersecurity requests augmented with two adversarial dialogue settings, and evaluate eight LLMs on it. Results show that prior assistant behavior dramatically changes responses to an unchanged request: among 400 request pairs, 376 are comparable, and compliance rises from 62.0% after a refusal history to 85.1% after an acceptance history. In the dialogue‑decomposition scenario the opposite trend appears; direct responses achieve 501/800 compliance, which drops to 172/800 after dialogue. Among 738 pairs that return model‑generated text in both conditions, the drop is 45.1 points. Failure feedback recovers only a small fraction of this loss.

Review

Original Source: https://arxiv.org/abs/2609.00578

[h] Back to Home