NeFut Logo NeFut
Admin Login

[CS.AI] MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models

Published at: 2026-09-07 22:00 Last updated: 2026-09-08 00:37
#AI #Machine Learning #Artificial Intelligence

As vision‑language models rapidly improve in image understanding, cross‑modal reasoning and complex instruction execution, the ability to follow instructions has become a primary measure of their reliability and practicality. Existing multimodal instruction‑following benchmarks, however, suffer from narrow language coverage and lack of adversarial safety scenarios, making them unsuitable for real‑world multilingual and safety‑critical applications. To fill this gap we introduce MM‑IFEval‑Pro, a benchmark that supports both Chinese and English tasks and incorporates diverse instruction‑hijacking cases. MM‑IFEval‑Pro comprises four major task categories, 24 sub‑categories, eight instruction types and 52 sub‑types; each sample carries on average 3.0 constraints to realistically emulate complex instruction settings. We also build a reinforcement‑learning training set enriched with Chinese and adversarial instructions. Experiments show that this training set markedly boosts model performance on MM‑IFEval‑Pro and transfers well to other mainstream multimodal benchmarks, demonstrating strong cross‑task and cross‑language generalization.

Review

Original Source: https://arxiv.org/abs/2609.04859

[h] Back to Home