NeFut Logo NeFut
Admin Login

[CS.AI] AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents

Published at: 2026-09-11 22:00 Last updated: 2026-09-12 06:35
#AI #Machine Learning #Open Source

This work introduces an end‑to‑end evaluation framework for image‑triggered command injection against computer‑use agents (CUAs). The framework tests whether a local visual patch can cause verifiable environmental effects across the full chain of screenshot input, vision‑language model (VLM) generation, action parsing, and environment execution. We train and deploy patches on author‑controlled GitHub Pages and a locally hosted CSDN clone, then evaluate them in real settings on five open‑source or publicly available GUI‑agent or VLM backends. Across 600 online instances the metrics T‑ASR, TAPR, and E2E‑ASR reach 84.5 %, 47.0 %, and 20.3 % respectively. Trajectory analysis reveals that in some successful cases the agent first runs a malicious terminal command and then resumes the original benign task. These findings indicate that optimized local visual signals can affect VLM outputs and propagate through the execution pipeline of open CUAs, posing real environmental risks.

Review

Original Source: https://arxiv.org/abs/2609.09212

[h] Back to Home