paper-with-me

Papers

HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding

2025-03-17 · Jiahe Zhao, Ruibing Hou, Zejie Tian, Hong Chang, Shiguang Shan

We propose a new task to benchmark human-in-scene understanding for embodied agents: Human-In-Scene Question Answering (HIS-QA). Given a human motion within a 3D scene, HIS-QA requires the agent to comprehend human states and behaviors, reason about its surrounding environment, and answer human-related questions within the scene. To support this new task, we present HIS-Bench, a multimodal benchmark that systematically evaluates HIS understanding across a broad spectrum, from basic perception to commonsense reasoning and planning. Our evaluation of various vision-language models on HIS-Bench reveals significant limitations in their ability to handle HIS-QA tasks. To this end, we propose HIS-GPT, the first foundation model for HIS understanding. HIS-GPT integrates 3D scene context and human motion dynamics into large language models while incorporating specialized mechanisms to capture human-scene interactions. Extensive experiments demonstrate that HIS-GPT sets a new state-of-the-art on HIS-QA tasks. We hope this work inspires future research on human behavior analysis in 3D scenes, advancing embodied AI and world models.

📄 PDF Abstract BibTeX arXiv:2503.12955

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringScene Understanding

Similar Papers 제목 키워드 기반

Evaluating Compositional Scene Understanding in Multimodal Generative Models

2025-03-29 · Shuhao Fu, Andrew Jun Lee, Anna Wang, Ida Momennejad 외

The visual world is fundamentally compositional. Visual scenes are defined by the composition of objects and their relations. Hence, it is essential for computer vision systems to reflect and exploit this compositionalit…

Scene Understanding

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

2025-01-25 · Jiaxing Zhao, Qize Yang, Yixing Peng, Detao Bai 외

In human-centric scenes, the ability to simultaneously understand visual and auditory information is crucial. While recent omni models can process multiple modalities, they generally lack effectiveness in human-centric s…

Action UnderstandingEmotion RecognitionLanguage ModelingLanguage Modelling+3

OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving

2025-09-24 · Pei Liu, Hongliang Lu, Haichao Liu, Haipeng Liu 외 arxiv

Human vision is capable of transforming two-dimensional observations into an egocentric three-dimensional scene understanding, which underpins the ability to translate complex scenes and exhibit adaptive behaviors. This …

Visual Question AnsweringKnowledge DistillationScene UnderstandingAutonomous Driving

Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning

2025-02-19 · Rui Zhao, Qirui Yuan, Jinyu Li, Haofeng Hu 외

End-to-end autonomous driving, which directly maps raw sensor inputs to low-level vehicle controls, is an important part of Embodied AI. Despite successes in applying Multimodal Large Language Models (MLLMs) for high-lev…

Autonomous DrivingBench2DriveMotion PlanningQuestion Answering+3

City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning

2025-07-17 · Penglei Sun, Yaoxian Song, Xiangru Zhu, Xiang Liu 외

Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have primarily focused on indoor household tasks, …

Question AnsweringScene Understanding