paper-with-me

홈 › Papers

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models

2026-05-05 · Kangkang Wang, Qinting Jiang, Wanping Zhang, Bowen Ren, Shengzhao Wen arxiv

Multidimensional human understanding is essential for real-world applications such as film analysis and virtual digital humans, yet current LVLM benchmarks largely focus on single-task settings and lack fine-grained, human-centric evaluation. In this work, we introduce MHPR, a comprehensive benchmark for joint perception-reasoning over human-centric scenes spanning individual, multi-person, and human-object interaction dimensions. MHPR comprises a multi-level data design-Captioned Raw Data (C-RD), Supervised Fine-Tuning Data (SFT-D), Reinforcement Learning Data (RL-D), and Test Data (T-D)-together with an automated caption/VQA generation pipeline (ACVG) that performs category-wise attribute decomposition, attribute-specific rewriting, and multi-model voting to ensure high-quality, scalable annotations. We evaluate state-of-the-art vision-language models on fine-grained attributes (appearance, clothing, pose, parts) and high-level semantics (social relations, action semantics, spatial relations, intent and functionality). Our findings show that: 1) format-aligned SFT data substantially improves instruction following and stability; 2) challenge-focused RL data derived from bad-case analysis further enhances perception and reasoning on difficult instances; and 3) training Qwen2.5-VL-7B with MHPR yields significant gains, achieving near-parity with considerably larger models. We release ACVG and MHPR to facilitate reproducible, extensible research on human-centric perception and reasoning.

📄 PDF Abstract BibTeX arXiv:2605.03485

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningInstruction Following

Similar Papers 제목 키워드 기반

MARVEL: Multidimensional Abstraction and Reasoning through Visual Evaluation and Learning

2024-04-21 · Yifan Jiang, Jiarui Zhang, Kexuan Sun, Zhivar Sourati 외

While multi-modal large language models (MLLMs) have shown significant progress on many popular visual reasoning benchmarks, whether they possess abstract visual reasoning abilities remains an open question. Similar to t…

Visual Reasoning

Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs

2025-05-28 · Aditya Kanade, Tanuja Ganu

Multimodal Large Language Models (MLLMs) show reasoning promise, yet their visual perception is a critical bottleneck. Strikingly, MLLMs can produce correct answers even while misinterpreting crucial visual elements, mas…

Multi-Dimensional, Nuanced and Subjective - Measuring the Perception of Facial Expressions

2022-01-01 · CVPR 2022 1 · De'Aira Bryant, Siqi Deng, Nashlie Sephus, Wei Xia 외

Humans can perceive multiple expressions, each one with varying intensity, in the picture of a face. We propose a methodology for collecting and modeling multidimensional modulated expression annotations from human a…

Facial Expression Recognition (FER)

A new visual quality metric for Evaluating the performance of multidimensional projections

2024-07-23 · Maniru Ibrahim, Thales Vieira

Multidimensional projections (MP) are among the most essential approaches in the visual analysis of multidimensional data. It transforms multidimensional data into two-dimensional representations that may be shown as sca…

Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmark

2026-04-09 · Longgang Zhang, Xiaowei Fu, Fuxiang Huang, Lei Zhang arxiv

Network traffic, as a key media format, is crucial for ensuring security and communications in modern internet infrastructure. While existing methods offer excellent performance, they face two key bottlenecks: (1) They f…

Multimodal Reasoning