NeFut Logo NeFut
中 Admin Login

[CS.AI] CART: Closed-Loop Adaptive Red Teaming for Large Language Models

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

CART (Closed-Loop Adaptive Red Teaming) is a framework that uses each test outcome to steer the next probe. It starts with broad risk coverage, follows emerging weaknesses, keeps new probes diverse, and records the evidence and source of every finding. The system separates three roles: the Challenger (generates tests), the Target (the model or a bounded tool‑using agent), and the Judge (evaluates results), allowing independent study of each component.

Across three evaluation families (Frontier, JAH, and Agentic), CART discovers more failures and raises average risk compared to static seed replay, outperforming every Target with an available baseline. The advantage also extends to tool‑mediated agent tests, indicating that contextual adaptation can expose weaknesses that direct prompt replay does not exercise. The results describe what the test policies can uncover, not how often failures occur in real deployments. Experiments also show that Challenger‑Judge pairings affect the evidence uncovered, highlighting the need for role separation and independent review.

Overall, CART turns red teaming from a one‑time checklist into a continuous, adaptive, and auditable search for model and agent weaknesses.

Review

Original Source: https://arxiv.org/abs/2609.27336

[h] Back to Home