paper-with-me

홈 › Papers

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding

2025-07-16 · Renjie Li, Ruijie Ye, Mingyang Wu, Hao Frank Yang, Zhiwen Fan, Hezhen Hu, Zhengzhong Tu arxiv

Humans are integral components of the transportation ecosystem, and understanding their behaviors is crucial to facilitating the development of safe driving systems. Although recent progress has explored various aspects of human behavior$\unicode{x2014}$such as motion, trajectories, and intention$\unicode{x2014}$a comprehensive benchmark for evaluating human behavior understanding in autonomous driving remains unavailable. In this work, we propose $\textbf{MMHU}$, a large-scale benchmark for human behavior analysis featuring rich annotations, such as human motion and trajectories, text description for human motions, human intention, and critical behavior labels relevant to driving safety. Our dataset encompasses 57k human motion clips and 1.73M frames gathered from diverse sources, including established driving datasets such as Waymo, in-the-wild videos from YouTube, and self-collected data. A human-in-the-loop annotation pipeline is developed to generate rich behavior captions. We provide a thorough dataset analysis and benchmark multiple tasks$\unicode{x2014}$ranging from motion prediction to motion generation and human behavior question answering$\unicode{x2014}$thereby offering a broad evaluation suite. Project page : https://MMHU-Benchmark.github.io.

📄 PDF Abstract BibTeX arXiv:2507.12463

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingQuestion Answering

Similar Papers 제목 키워드 기반

Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning

2024-06-15 · Jifan Zhang, Lalit Jain, Yang Guo, Jiayi Chen 외

We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2 million captions, collected through crowdsourcing rating data for The New Yorker's weekly…

Caption Generation

Borrowing Human Senses: Comment-Aware Self-Training for Social Media Multimodal Classification

2023-03-27 · Chunpu Xu, Jing Li

Social media is daily creating massive multimedia content with paired image and text, presenting the pressing need to automate the vision and language understanding for various multimodal classification tasks. Compared t…

ClassificationHate Speech DetectionRelation ClassificationSarcasm Detection+2

Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset

2022-05-25 · Ashish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, Radu Soricut

Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets. In this paper we present the Crossmodal-3600 dataset (XM3600 in short), a geographically diver…

Image CaptioningImage RetrievalImage-text RetrievalImage-to-Text Retrieval+3

AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification

2026-05-07 · Claire Chen, Jiabao Sean Xiao, Shuze Daniel Liu, Facundo Perez Paolino 외 arxiv

Modern astronomical observatories generate a massive volume of multimodal data, creating a critical bottleneck for expert human review. While multimodal large language models (LLMs) have shown promise in interpreting com…

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

2023-11-27 · CVPR 2024 1 · Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng 외

We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected m…

Complex Query AnsweringLogical ReasoningVisual Reasoning