paper-with-me

홈 › Papers

VLH: Vision-Language-Haptics Foundation Model

2025-08-02 · Luis Francisco Moreno Fuentes, Muhammad Haris Khan, Miguel Altamirano Cabrera, Valerii Serpiva, Dmitri Iarchuk, Yara Mahmoud, Issatay Tokmurziyev, Dzmitry Tsetserukou arxiv

We present VLH, a novel Visual-Language-Haptic Foundation Model that unifies perception, language, and tactile feedback in aerial robotics and virtual reality. Unlike prior work that treats haptics as a secondary, reactive channel, VLH synthesizes mid-air force and vibration cues as a direct consequence of contextual visual understanding and natural language commands. Our platform comprises an 8-inch quadcopter equipped with dual inverse five-bar linkage arrays for localized haptic actuation, an egocentric VR camera, and an exocentric top-down view. Visual inputs and language instructions are processed by a fine-tuned OpenVLA backbone - adapted via LoRA on a bespoke dataset of 450 multimodal scenarios - to output a 7-dimensional action vector (Vx, Vy, Vz, Hx, Hy, Hz, Hv). INT8 quantization and a high-performance server ensure real-time operation at 4-5 Hz. In human-robot interaction experiments (90 flights), VLH achieved a 56.7% success rate for target acquisition (mean reach time 21.3 s, pose error 0.24 m) and 100% accuracy in texture discrimination. Generalization tests yielded 70.0% (visual), 54.4% (motion), 40.0% (physical), and 35.0% (semantic) performance on novel tasks. These results demonstrate VLH's ability to co-evolve haptic feedback with perceptual reasoning and intent, advancing expressive, immersive human-robot interactions.

📄 PDF Abstract BibTeX arXiv:2508.01361

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Gesture-Based Visual Learning Model for Acoustophoretic Interactions using a Swarm of AcoustoBots

2026-04-21 · Alex Lin, Lei Gao, Narsimlu Kemsaram, Sriram Subramanian arxiv

AcoustoBots are mobile acoustophoretic robots capable of delivering mid-air haptics, directional audio, and acoustic levitation, but existing implementations rely on scripted commands and lack an intuitive interface for …

Robust Robotic Pouring using Audition and Haptics

2020-02-29 · Hongzhuo Liang, Chuangchuang Zhou, Shuang Li, Xiaojian Ma 외

Robust and accurate estimation of liquid height lies as an essential part of pouring tasks for service robots. However, vision-based methods often fail in occluded conditions while audio-based methods cannot work well in…

Co-reference via Pointing and Haptics in Multi-Modal Dialogues

2012-06-01 · NAACL 2012 6 · Lin Chen, Barbara Di Eugenio

Guiding Interaction Behaviors for Multi-modal Grounded Language Learning

2017-08-01 · WS 2017 8 · Jesse Thomason, Jivko Sinapov, Raymond Mooney

Multi-modal grounded language learning connects language predicates to physical properties of objects in the world. Sensing with multiple modalities, such as audio, haptics, and visual colors and shapes while performing …

Grounded language learningRetrieval

ETHOS: A Robotic Encountered-Type Haptic Display for Social Interaction in Virtual Reality

2025-11-07 · Eric Godden, Jacquie Groenewegen, Matthew K. X. J. Pan arxiv

We present ETHOS (Encountered-Type Haptics for On-demand Social Interaction), a dynamic encountered-type haptic display (ETHD) that enables natural physical contact in virtual reality (VR) during social interactions such…