As vision‑language models rapidly improve in image understanding, cross‑modal reasoning and complex instruction execution, the ability to follow instructions has become a primary measure of their reliability and practicality. Existing multimodal instruction‑following benchmarks, however, suffer from narrow language coverage and lack of adversarial safety scenarios, making them unsuitable for real‑world multilingual and safety‑critical applications. To fill this gap we introduce MM‑IFEval‑Pro, a benchmark that supports both Chinese and English tasks and incorporates diverse instruction‑hijacking cases. MM‑IFEval‑Pro comprises four major task categories, 24 sub‑categories, eight instruction types and 52 sub‑types; each sample carries on average 3.0 constraints to realistically emulate complex instruction settings. We also build a reinforcement‑learning training set enriched with Chinese and adversarial instructions. Experiments show that this training set markedly boosts model performance on MM‑IFEval‑Pro and transfers well to other mainstream multimodal benchmarks, demonstrating strong cross‑task and cross‑language generalization.
Review