Contextual security defenses synthesize a task‑specific policy and enforce it on an AI agent’s tool calls, preventing rogue actions.
In multi‑step tasks, the validity of an action often depends on what the agent has already done and learned. To address this, we introduce Sapien, a policy engine that enforces stateful contextual policies.
A Sapien policy describes permitted tool‑call sequences with an extended regular expression that incorporates
- Stateful predicates: dynamically decide future allowances based on the agent’s history;
- Deferred policy generation: refine constraints at runtime using the latest context;
- Scoped semantic checks: validate calls within specific semantic scopes.
Empirical results show that Sapien stays within a few percent of an unconstrained agent’s utility while ruling out 93%‑95% of attacks on AgentDojo and 62%‑85% on Toolathlon—roughly twice the protection offered by simple tool allowlists on long‑horizon tasks.
Review: Sapien embeds state information directly into policy expressions, enabling fine‑grained control over complex multi‑step interactions. It achieves a strong trade‑off between security and utility, highlighting the promise of scalable contextual security frameworks.