NeFut Logo NeFut
Admin Login

[CS.AI] ASCII Attack: Recontextualizing Harmful Requests as Artistic Critique in Large Language Models

Published at: 2026-09-03 22:00 Last updated: 2026-09-04 02:14
#AI #Machine Learning #LLM

Safety alignment usually trains large models to refuse overt harmful requests, but the training focuses on surface form. When the same operational content is recontextualized, the model’s defenses weaken. The ASCII Attack is such a recontextualization. In a single‑turn black‑box interaction it sends one message: the full harmful request is rendered in ASCII‑art characters as an “artwork” and the model is asked for artistic feedback. The request remains readable and is not hidden. The model’s reply is framed as artistic critique and may contain operational details that a plain request would have been denied.

In the experiments each framed prompt is paired with a direct‑question control to isolate the effect of surface wording while keeping topic, model and decoding constant. Across eleven models and eight harm topics, a harm‑aware classifier judges 62% of framed prompts as harmful versus 42% of controls. On the most vulnerable model the framed prompt succeeds 93% of the time. A single query matches or exceeds published single‑query attacks on four of five harm judges. The pattern shows that harm likelihood tracks the model more than the topic and does not diminish with scale. Nearly two‑thirds of framed rows have at least one judge dissenting from the majority, a finding that highlights measurement validity issues and suggests mismatched generalization.

Review

Original Source: https://arxiv.org/abs/2609.02215

[h] Back to Home