paper-with-me

홈 › Papers

Proactive Human-Robot Interaction using Visuo-Lingual Transformers

2023-10-04 · Pranay Mathur

Humans possess the innate ability to extract latent visuo-lingual cues to infer context through human interaction. During collaboration, this enables proactive prediction of the underlying intention of a series of tasks. In contrast, robotic agents collaborating with humans naively follow elementary instructions to complete tasks or use specific hand-crafted triggers to initiate proactive collaboration when working towards the completion of a goal. Endowing such robots with the ability to reason about the end goal and proactively suggest intermediate tasks will engender a much more intuitive method for human-robot collaboration. To this end, we propose a learning-based method that uses visual cues from the scene, lingual commands from a user and knowledge of prior object-object interaction to identify and proactively predict the underlying goal the user intends to achieve. Specifically, we propose ViLing-MMT, a vision-language multimodal transformer-based architecture that captures inter and intra-modal dependencies to provide accurate scene descriptions and proactively suggest tasks where applicable. We evaluate our proposed model in simulation and real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2310.02506

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exploring Human-Robot Collaboration: Analysis of Interaction Modalities in Challenging Tasks

2026-05-13 · Simone Arreghini, Cristina Iani, Alessandro Giusti, Valeria Villani 외 arxiv

This work compares three interaction modalities for human-robot collaboration: passive, reactive, and proactive. We studied 18 participants assembling a seven-layer colored tower from memory while using nearby and distan…

Commonsense Scene Semantics for Cognitive Robotics: Towards Grounding Embodied Visuo-Locomotive Interactions

2017-09-15 · Jakob Suchan, Mehul Bhatt

We present a commonsense, qualitative model for the semantic grounding of embodied visuo-spatial and locomotive interactions. The key contribution is an integrative methodology combining low-level visual processing with …

Learning Predictive Visuomotor Coordination

2025-03-30 · Wenqi Jia, Bolin Lai, Miao Liu, Danfei Xu 외

Understanding and predicting human visuomotor coordination is crucial for applications in robotics, human-computer interaction, and assistive technologies. This work introduces a forecasting-based task for visuomotor mod…

When May I Help You? On The Effect of Proactivity on Group Human-Robot Collaboration

2026-06-26 · Thomas Vitry, Vanessa Maeder, Kieran von Valeburg, Asihati Hazaiti 외 arxiv

Robot initiative is a central challenge in multi-party human-robot collaboration. A robot that contributes without being addressed may provide timely support, but it may also disrupt coordination, divide attention, or in…

The functional and temporal roles of gaze evolve across the phases and constraints of multi-stage robot-mediated manipulation

2026-06-20 · Manuela Uliano, Silvia Fattorini, Marco Controzzi arxiv

Goal-directed eye movements are a fundamental component of visuomotor control, enabling humans to anticipate and guide their actions. For this reason, they are increasingly used in human-robot interaction to estimate use…