paper-with-me

홈 › Papers

HumanVLM: Foundation for Human-Scene Vision-Language Model

2024-11-05 · Dawei Dai, Xu Long, Li Yutang, Zhang YuanHui, Shuyin Xia

Human-scene vision-language tasks are increasingly prevalent in diverse social applications, yet recent advancements predominantly rely on models specifically tailored to individual tasks. Emerging research indicates that large vision-language models (VLMs) can enhance performance across various downstream vision-language understanding tasks. However, general-domain models often underperform in specialized fields. This study introduces a domain-specific Large Vision-Language Model, Human-Scene Vision-Language Model (HumanVLM), designed to provide a foundation for human-scene Vision-Language tasks. Specifically, (1) we create a large-scale human-scene multimodal image-text dataset (HumanCaption-10M) sourced from the Internet to facilitate domain-specific alignment; (2) develop a captioning approach for human-centered images, capturing human faces, bodies, and backgrounds, and construct a high-quality Human-Scene image-text dataset (HumanCaptionHQ, about 311k pairs) that contain as much detailed information as possible about human; (3) Using HumanCaption-10M and HumanCaptionHQ, we train a HumanVLM. In the experiments, we then evaluate our HumanVLM across varous downstream tasks, where it demonstrates superior overall performance among multimodal models of comparable scale, particularly excelling in human-related tasks and significantly outperforming similar models, including Qwen2VL and ChatGPT-4o. HumanVLM, alongside the data introduced, will stimulate the research in human-around fields.

📄 PDF Abstract BibTeX arXiv:2411.03034

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingmodel

Similar Papers 제목 키워드 기반

doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation

2024-12-08 · Parthib Roy, Srinivasa Perisetla, Shashank Shriram, Harsha Krishnaswamy 외

Human-interactive robotic systems, particularly autonomous vehicles (AVs), must effectively integrate human instructions into their motion planning. This paper introduces doScenes, a novel dataset designed to facilitate …

Autonomous DrivingAutonomous VehiclesMotion PlanningVision-Language Navigation

CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes

2026-08-19 · Kumal Hewagamage, Isuranga Senavirathne, Sasika Amarasinghe, Hasitha Gallella 외 arxiv

4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vision encoders are largely limited to static 2D images or 3D point clouds w…

Contrastive LearningPoint Clouds

HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding

2025-03-17 · Jiahe Zhao, Ruibing Hou, Zejie Tian, Hong Chang 외

We propose a new task to benchmark human-in-scene understanding for embodied agents: Human-In-Scene Question Answering (HIS-QA). Given a human motion within a 3D scene, HIS-QA requires the agent to comprehend human state…

Question AnsweringScene Understanding

InHabit: Leveraging Image Foundation Models for Scalable 3D Human Placement

2026-04-21 · Nikita Kister, Pradyumna YM, István Sárándi, Jiayi Wang 외 arxiv

Training embodied agents to understand 3D scenes as humans do requires large-scale data of people meaningfully interacting with diverse environments, yet such data is scarce. Real-world capture is costly and limited to c…

Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding

2024-09-05 · Yunze Man, Shuhong Zheng, Zhipeng Bao, Martial Hebert 외

Complex 3D scene understanding has gained increasing attention, with scene encoding strategies playing a crucial role in this success. However, the optimal scene encoding strategies for various scenarios remain unclear, …

Question AnsweringScene UnderstandingVisual Grounding