paper-with-me

홈 › Papers

LEGS-POMDP: Language and Gesture-Guided Object Search in Partially Observable Environments

2026-03-05 · Ivy Xiao He, Stefanie Tellex, Jason Xinyu Liu arxiv

To assist humans in open-world environments, robots must interpret ambiguous instructions to locate desired objects. Foundation model-based approaches excel at multimodal grounding, but they lack a principled mechanism for modeling uncertainty in long-horizon tasks. In contrast, Partially Observable Markov Decision Processes (POMDPs) provide a systematic framework for planning under uncertainty but are often limited in supported modalities and rely on restrictive environment assumptions. We introduce LanguagE and Gesture-Guided Object Search in Partially Observable Environments (LEGS-POMDP), a modular POMDP system that integrates language, gesture, and visual observations for open-world object search. Unlike prior work, LEGS-POMDP explicitly models two sources of partial observability: uncertainty over the target object's identity and its spatial location. In simulation, multimodal fusion significantly outperforms unimodal baselines, achieving an average success rate of 89\% across challenging environments and object categories. Finally, we demonstrate the full system on a quadruped mobile manipulator, where real-world experiments qualitatively validate robust multimodal perception and uncertainty reduction under ambiguous instructions.

📄 PDF Abstract BibTeX arXiv:2603.04705

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling

2025-01-31 · Pinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu 외

Generating full-body human gestures based on speech signals remains challenges on quality and speed. Existing approaches model different body regions such as body, legs and hands separately, which fail to capture the spa…

DenoisingGesture Generation

Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation

2026-03-18 · Rui Hong, Jana Kosecka arxiv

Estimating 3D hand pose from monocular RGB images is fundamental for applications in AR/VR, human-computer interaction, and sign language understanding. In this work we focus on a scenario where a discrete set of gesture…

3D Hand Pose Estimation3D Pose Estimation

LEGS: Fine-Tuning Teleop-Free VLAs for Humanoid Loco-manipulation in an Embodied Gaussian Splatting World

2026-05-31 · Hojune Kim, Timothy Chen, Jiankai Sun, Lars W. Osterberg 외 arxiv

Training vision-language-action (VLA) policies for humanoid loco-manipulation is constrained by the high cost and complexity of collecting human teleoperation demonstrations. VLA policies fine-tuned in simulators have, u…

SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where

2025-09-28 · Yiheng Huang, Junran Peng, Silei Shen, Jingwei Yang 외 arxiv

The accompanying actions and gestures in dialogue are often closely linked to interactions with the environment, such as looking toward the interlocutor or using gestures to point to the described target at appropriate m…

Gesture Generation

LEGS: Laplacian-Enhanced Gaussian Splatting with a Nonlinear Weighted Loss

2026-06-06 · Yongfei Guo, Qizhou Huo, Xuan Sun, Yuanhao Gong arxiv

3D Gaussian Splatting (3DGS) has become an efficient explicit representation for radiance field reconstruction and real-time novel view synthesis. However, its standard photometric loss treats flat and structure-rich reg…

Novel View Synthesis