paper-with-me

홈 › Papers

MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models

2026-09-04 · Changming Xiao, Zhenliang Ni, Jinhui He, Han Shu, Jie Hu arxiv

As vision-language models (VLMs) rapidly advance in image understanding, cross-modal reasoning, and complex instruction execution, instruction-following capability has become a key indicator of their reliability and practicality. However, existing multimodal instruction-following benchmarks still suffer from limited language coverage and insufficient adversarial safety scenarios, making them inadequate for evaluating real-world multilingual and safety-sensitive settings. To address these gaps, we present MM-IFEval-Pro, a multimodal instruction-following benchmark covering Chinese and English tasks as well as diverse instruction hijacking cases. MM-IFEval-Pro includes 4 major task categories and 24 subcategories and 8 instruction categories with 52 subcategories, with each sample containing an average of 3.0 constraints to realistically simulate complex instruction scenarios. We further construct a reinforcement-learning training set enriched with Chinese and adversarial instructions, which significantly improves model performance on MM-IFEval-Pro and transfers effectively to other mainstream multimodal benchmarks, demonstrating strong cross-task and cross-language generalization.

📄 PDF Abstract BibTeX arXiv:2609.04859

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

M-IFEval: Multilingual Instruction-Following Evaluation

2025-02-07 · Antoine Dussolle, Andrea Cardeña Díaz, Shota Sato, Peter Devine

Instruction following is a core capability of modern Large language models (LLMs), making evaluating this capability essential to understanding these models. The Instruction Following Evaluation (IFEval) benchmark from t…

Instruction Following

IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages

2026-02-25 · Thanmay Jayakumar, Mohammed Safi Ur Rahman Khan, Raj Dabre, Ratish Puduppully 외 arxiv

Instruction-following benchmarks remain predominantly English-centric, leaving a critical evaluation gap for the hundreds of millions of Indic language speakers. We introduce IndicIFEval, a benchmark evaluating constrain…

IFEvalCode: Controlled Code Generation

2025-07-30 · Jian Yang, Wei Zhang, Shukai Liu, Linzheng Chai 외 arxiv

Code large language models (Code LLMs) have made significant progress in code generation by translating natural language descriptions into functional code; however, real-world applications often demand stricter adherence…

Code Generation

Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models

2025-07-16 · Bo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng 외 arxiv

Instruction-following capability has become a major ability to be evaluated for Large Language Models (LLMs). However, existing datasets, such as IFEval, are either predominantly monolingual and centered on English or si…

Instruction Following

EuroAlpaca: Task-Preserving Localisation of Instruction Data for European Languages

2026-09-04 · Aleix Sant, Jordi Luque, Carlos Escolano arxiv

Machine translation (MT) offers a scalable way to extend English instruction-tuning data to multiple languages, but it can distort task-critical constraints and required outputs, creating corrupted training examples and …

Instruction FollowingMachine Translation