paper-with-me

홈 › Papers

HCSG: Human-Centric Semantic-Geometric Reasoning for Vision-Language Navigation

2026-05-13 · Haoxuan Xu, Tianfu Li, Wenbo Chen, Yi Liu, Jin Wu, Huashuo Lei, Yunfan Lou, Lujia Wang, Hesheng Wang, Haoang Li arxiv

VLN has achieved remarkable progress by scaling data and model capacity. However, the assumption of a static environment breaks down in real-world indoor scenarios, where robots inevitably encounter dynamic pedestrians. Existing human-aware approaches typically treat humans merely as moving obstacles based on implicit visual cues, lacking the explicit reasoning required to interpret human intentions or maintain social norms. To address this, we propose HCSG, the first human-centric framework for VLN. This framework provides a robust foundation for safe, socially intelligent navigation in dynamic human-robot environments that shifts the paradigm from passive collision avoidance to active human behavior understanding. Specifically, HCSG introduces a unified Human Understanding Module that synergizes two key capabilities: (i) geometric forecasting, which predicts human pose and trajectory to anticipate future motion dynamics; and (ii) semantic interpretation, which leverages a Vision-Language Model (VLM) to generate natural language descriptions of human actions and intentions. These semantic-geometric representations are fused into the agent's topological map for instruction-conditioned planning. Furthermore, a social distance loss is introduced to enforce socially compliant interaction distances. Extensive experiments on the HA-VLNCE benchmark demonstrate that HCSG significantly outperforms state-of-the-art methods, achieving a 14% improvement in Success Rate and a 34% reduction in Collision Rate. Our project can be seen at https://haoxuanxu1024.github.io/HCSG/.

📄 PDF Abstract BibTeX arXiv:2605.13321

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language NavigationCollision Avoidance

Similar Papers 제목 키워드 기반

Allocentric Perceiver: Disentangling Allocentric Reasoning from Egocentric Visual Priors via Frame Instantiation

2026-02-05 · Hengyi Wang, Ruiqiang Zhang, Chang Liu, Guanjie Wang 외 arxiv

With the rising need for spatially grounded tasks such as Vision-Language Navigation/Action, allocentric perception capabilities in Vision-Language Models (VLMs) are receiving growing focus. However, VLMs remain brittle …

Vision-Language NavigationSpatial Reasoning

Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning

2025-09-24 · Xun Li, Rodrigo Santa Cruz, Mingze Xi, Hu Zhang 외 arxiv

To enable robots to comprehend high-level human instructions and perform complex tasks, a key challenge lies in achieving comprehensive scene understanding: interpreting and interacting with the 3D environment in a meani…

Scene UnderstandingPoint Clouds

Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention

2025-12-30 · Haijing Liu, Zhiyuan Song, Hefeng Wu, Tao Pu 외 arxiv

Egocentric Referring Video Object Segmentation (Ego-RVOS) aims to segment the specific object actively involved in a human action, as described by a language query, within first-person videos. This task is critical for u…

Referring Video Object Segmentation

HumanMoveVQA: Can Video MLLMs reason about human movement in videos?

2026-06-26 · Pulkit Gera, Faegheh Sardari, Asmar Nadeem, Valentina Bono 외 arxiv

Despite the rapid advance of Multimodal Large Language Models (MLLMs) in high-level video understanding, a fundamental bottleneck remains: these models collapse complex human motion into coarse semantic labels. Existing …

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment

2025-08-29 · Jinzhou Tang, Jusheng zhang, Sidi Liu, Waikit Xiu 외 arxiv

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene u…

Scene UnderstandingQuestion Answering