NeFut Logo NeFut
Admin Login

[CS.AI] Adversarial Pragmatics: A New Benchmark for AI Safety Evaluation

Published at: 2026-07-19 22:00 Last updated: 2026-07-22 01:02
#AI #Machine Learning #Open Source

Abstract

Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments.

This paper introduces adversarial pragmatics as a benchmark and annotation protocol for evaluating model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, indirect speech acts, and multi-turn agent transcripts. The contribution is empirical and methodological:

The benchmark treats labels as inference licenses: it tests whether safety-relevant categories project across paraphrase, wrapper, model, and judge condition. In the pilot, a rubric-aided LLM judge graded its own outputs with expected-behaviour fields visible and still missed the safety-relevant minority classes.

Blogger's Review: This paper offers a new perspective on the safety evaluation of AI models, revealing their performance in complex linguistic contexts through systematic methodologies and comprehensive benchmarks. The introduction of adversarial pragmatics will aid in better understanding and improving model behavior, especially when handling ambiguous and complex instructions.

Original Source: https://arxiv.org/abs/2607.01153

[h] Back to Home