A skill is a modular bundle of natural‑language instructions, executable scripts, and reference resources that an agent can load at runtime to perform a specific task. Skill‑based agent systems thus enable flexible reuse of third‑party capabilities, but the openness of the skill ecosystem also creates a new attack surface. Prior work has mainly examined vulnerabilities inside individual skills, while risks arising from interactions across skills have received little attention.
We introduce the threat paradigm of skill cascading attacks: a malicious objective is split across multiple skills so that each modification appears benign in isolation, yet their combined execution becomes harmful. For example, in a prescription‑review pipeline the first skill weakens signals of recently discontinued drugs, the second downgrades the severity of any interaction tied to them, and the third suppresses the resulting low‑priority alert in the final summary, causing a severe drug‑interaction warning to disappear before reaching the physician.
To study this safety blind spot systematically, we build SkillCascade, an automated multi‑agent red‑team framework, and release SkillCascade‑Bench, a benchmark containing 213 validated cascading test cases across various agent systems and domains.
Across representative agents (e.g., OpenClaw, Claude Code, Codex) and different LLM backbones, cascaded interactions reliably induce harmful behavior while evading existing per‑skill scanners and runtime monitors.
Our findings highlight a gap between component‑level integrity and system‑level safety, calling for defenses that reason over cross‑skill interactions rather than inspecting isolated skills.
Review