Abstract
Foundation model safety benchmarks capture the AI risks at the time of their publication. As models improve and governments enact new AI safety legislation, these risk taxonomies become incomplete and attack prompts ineffective. We present AIR-BENCH Live, a self-evolving successor to AIR-BENCH 2024.
An automated update pipeline monitors government regulations and classifies new policies against the current four-tier risk taxonomy, either matching them to existing categories or proposing new granular categories. A multi-agent, persona-driven prompt generation algorithm generates realistic, multilingual prompts with minimal human review, allowing room for improvement with modern jailbreaking techniques. This algorithm overhauls legacy prompts and generates prompts for new categories.
In our current version, the benchmark has expanded from 314 to 335 granular risks, with 21 new categories drawn from 31 truly novel policy clauses across seven jurisdictions. Evaluating 14 recent models, we find a wide safety spread (from 0.17 to 0.89 among the models judged on their own behavior). The modernized prompts are on average 0.06 points harder than the 2024 set, with the largest drops concentrated among the most compliant models, and most models are modestly less safe on non-English prompts. By continuously absorbing new regulation and regenerating prompts, AIR-BENCH Live is designed to evolve alongside a fast-moving field.
Blogger's Review: The innovation of AIR-BENCH Live lies in its automated update capability, allowing continuous adaptation to changing AI safety regulations. This self-evolving benchmark not only enhances the accuracy of safety assessments but also provides new directions for model improvements, reflecting the close interplay between AI governance and technological advancement. Balancing safety and innovation will be a critical challenge moving forward.