NeFut Logo NeFut
中 Admin Login

[CS.AI] Ajar: Measuring Open Privilege in Agent Defenses

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#algorithm #AI #Machine Learning

A language‑model agent operates through the tools it is given. The data it reads while solving a task can steer how it uses those tools. Consequently, many safe‑execution techniques sit between the agent and its tools, enforcing access control, information‑flow or isolation at that boundary. Current agent‑security benchmarks evaluate defenses mainly via indirect prompt injection, measuring how much they reduce attack success while preserving the agent’s utility. A defense is judged only on the agent’s execution, so it may score well on both metrics while still leaving open a transfer, deletion or broad read that no task actually requires.

Ajar measures this “open privilege” directly. It attaches to an existing agent‑security benchmark, reusing the benchmark’s tasks, tool schemas, reference solutions and goal states. For each benign task, Ajar generates candidate tool calls that the task does not need, and presents these calls to the defense at every point where the agent could act. Allowing any of these calls indicates privilege left open.

We integrate Ajar with AgentDojo, turning open privilege into a third evaluation axis alongside attack success and benign utility. We run it on five defenses: Progent, CaMeL, AC4A, Permission Assistant, and Claude Code’s Auto mode. The defenses leave vastly different amounts of privilege open; two defenses leak almost the same amount yet differ widely in the number of benign tasks they complete; another tightens its profile by refusing calls that the tasks were entitled to make. The amount of open privilege cannot be inferred from measured attack success or utility.

The source code of Ajar is available at https://github.com/reSHARMA/Ajar.

Review

Original Source: https://arxiv.org/abs/2609.26900

[h] Back to Home