NeFut Logo NeFut
中 Admin Login

[CS.AI] MLCommons Jailbreak Benchmark v1.0

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Modern AI systems are designed to refuse hazardous requests. A jailbreak prompt is a specially crafted instruction that bypasses safety safeguards and elicits outputs the model would normally refuse. The MLCommons Jailbreak Benchmark v1.0 proposes an end‑to‑end methodology for evaluating the robustness of large language models against single‑turn text‑based jailbreak attacks. The pipeline includes criteria‑driven selection of systems and attacks, paired baseline and adversarial evaluation, human annotation, automated evaluator calibration, scoring, grading, and risk‑calibrated disclosure.

The benchmark evaluates eight open‑weight models using 264 seed prompts that span eleven hazard categories, with representative attacks drawn from the MLCommons Jailbreak Taxonomy. Responses are assessed with the AILuminate Assessment Standard v1.4, and robustness is measured by the “Resilience Gap,” the change in safety performance between baseline and adversarial conditions.

Results show that the unsafe‑response rate rises from 11.08% under baseline to 18.65% under jailbreak, yielding an average Resilience Gap of 7.57%. Accessible models exhibit a larger mean gap, and effectiveness varies markedly across attack categories and hazards. The benchmark also examines evaluator reliability and sources of measurement error.

Beyond reporting numbers, Jailbreak Benchmark v1.0 establishes a reproducible methodological foundation for comparative jailbreak evaluation and enables future expansion across models, attacks, hazards, and evaluation methods.

Review

Original Source: https://arxiv.org/abs/2610.02827

[h] Back to Home