paper-with-me

Papers

HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding

2025-07-07 · Yuxuan Cai, Jiangning Zhang, Zhenye Gan, Qingdong He, Xiaobin Hu, Junwei Zhu, Yabiao Wang, Chengjie Wang, Zhucun Xue, Chaoyou Fu, Xinwei He, Xiang Bai arxiv

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks involving both images and videos. However, their capacity to comprehend human-centric video data remains underexplored, primarily due to the absence of comprehensive and high-quality evaluation benchmarks. Existing human-centric benchmarks predominantly emphasize video generation quality and action recognition, while overlooking essential perceptual and cognitive abilities required in human-centered scenarios. Furthermore, they are often limited by single-question paradigms and overly simplistic evaluation metrics. To address above limitations, we propose a modern HV-MMBench, a rigorously curated benchmark designed to provide a more holistic evaluation of MLLMs in human-centric video understanding. Compared to existing human-centric video benchmarks, our work offers the following key features: (1) Diverse evaluation dimensions: HV-MMBench encompasses 13 tasks, ranging from basic attribute perception (e.g., age estimation, emotion recognition) to advanced cognitive reasoning (e.g., social relationship prediction, intention prediction), enabling comprehensive assessment of model capabilities; (2) Varied data types: The benchmark includes multiple-choice, fill-in-blank, true/false, and open-ended question formats, combined with diverse evaluation metrics, to more accurately and robustly reflect model performance; (3) Multi-domain video coverage: The benchmark spans 50 distinct visual scenarios, enabling comprehensive evaluation across fine-grained scene variations; (4) Temporal coverage: The benchmark covers videos from short-term (10 seconds) to long-term (up to 30min) durations, supporting systematic analysis of models temporal reasoning abilities across diverse contextual lengths.

📄 PDF Abstract BibTeX arXiv:2507.04909

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionAction RecognitionVideo GenerationAge Estimation

Similar Papers 제목 키워드 기반

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

2025-08-18 · Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to…

EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos

2025-03-28 · YuXuan Li, Vijay Veerabadran, Michael L. Iuzzolino, Brett D. Roads 외

We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice video QA instances for the Ego4D datase…

BenchmarkingQuestion AnsweringVideo Question Answering

In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting

2025-09-09 · Taiying Peng, Jiacheng Hua, Miao Liu, Feng Lu arxiv

The emergence of advanced multimodal large language models (MLLMs) has significantly enhanced AI assistants' ability to process complex information across modalities. Recently, egocentric videos, by directly capturing us…

Video Question AnsweringGaze Estimation

Ego-Grounding for Personalized Question-Answering in Egocentric Videos

2026-04-02 · Junbin Xiao, Shenglang Zhang, Pengxiang Zhu, Angela Yao arxiv

We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this …

HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding

2024-10-09 · Keliang Li, Zaifei Yang, Jiahe Zhao, Hongze Shen 외

The significant advancements in visual understanding and instruction following from Multimodal Large Language Models (MLLMs) have opened up more possibilities for broader applications in diverse and universal human-centr…

BenchmarkingInstruction Following