paper-with-me

홈 › Papers

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

2026-06-03 · Huangchen Xu, Yuan Wu, Yi Chang arxiv

Multimodal large language models have made rapid progress in video understanding, yet existing benchmarks largely rely on simple prompts and provide limited evidence about whether models can satisfy explicit output constraints. We introduce VCIFBench, a benchmark for evaluating complex instruction following in video understanding. VCIFBench constructs constraint-rich instructions from both benchmark-adapted and directly video-grounded prompts, covering content, format, style, and structure requirements, and evaluates model outputs with a hybrid verification pipeline. The benchmark contains 306 satisfiable test instructions, a 540-pair DPO preference dataset, and a 30-item conflict diagnostic subset. Experiments on 10 MLLMs show that joint constraint satisfaction remains challenging. We further show that DPO training on VCIFBench data can improve instruction-following performance.

📄 PDF Abstract BibTeX arXiv:2606.04588

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

2026-08-26 · Hongbo Liu, Peixian Chen, Sihan Liu, Peiyuan Zhang 외 arxiv

Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires …

Instruction Following

IF-VidCap: Can Video Caption Models Follow Instructions?

2025-10-21 · Shihao Li, Yuanxing Zhang, Jiangtao Wu, Zhide Lei 외 arxiv

Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, uncon…

Video CaptioningDense Captioning

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

2026-06-01 · Huiqiong Li, Jiayu Wang, Zhiting Mei, Anirudha Majumdar 외 arxiv

Video world models are increasingly used in robotic manipulation, yet existing benchmarks mostly evaluate them under valid, feasible, and safe instructions. We introduce RoboTrustBench, a benchmark for evaluating the tru…

Instruction Following

Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation

2026-07-07 · Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian 외 arxiv

Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task for NLPCC 2026. DA-MIVQA extends previous …

Video Question AnsweringVisual Grounding

InFoBench: Evaluating Instruction Following Ability in Large Language Models

2024-01-07 · Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao 외

This paper introduces the Decomposed Requirements Following Ratio (DRFR), a new metric for evaluating Large Language Models' (LLMs) ability to follow instructions. Addressing a gap in current methodologies, DRFR breaks d…

Instruction Following