NeFut Logo NeFut
Admin Login

[CS.AI] Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation

Published at: 2026-08-29 22:00 Last updated: 2026-08-30 12:07
#AI #Machine Learning

Audio‑video generation is shifting from prompt‑only synthesis to multimodal conditioning, where text, images, audio and video jointly shape the output. This shift moves safety evaluation from detecting harmful intent in a single input to spotting risks that emerge from cross‑modal and temporal interactions. Existing safety benchmarks remain largely prompt‑centric or tied to fixed conditioning interfaces, making systematic study of such compositional risks difficult.

To address this gap we introduce Multi2AV‑Safety, the first benchmark that covers all 11 non‑singleton T/I/A/V conditioning configurations for audio‑video generation, containing 11,024 attack instances. Evaluation shows systematic weaknesses in representative multimodal safety guards across attack mechanisms and harm‑evidence structures.

Two complementary failure modes are observed: harmful semantics can arise from the combination of individually benign inputs, and explicit harmful cues become harder to detect when mixed with benign multimodal context.

These findings highlight "compositional risk perception" as a central capability gap: current guards fail to reliably integrate safety evidence across modalities and time even when all conditioning inputs are observable. The dataset will be publicly released in October 2026.

Blogger's Review: Multi2AV‑Safety exposes a critical blind spot in safety assessment for multimodal generation and offers a concrete roadmap for building more robust protection mechanisms.

Original Source: https://arxiv.org/abs/2608.26535

[h] Back to Home