paper-with-me

Papers

RoboOmni: Proactive Robot Manipulation in Omni-modal Context

2025-10-27 · Siyin Wang, Jinlan Fu, Feihong Liu, Xinzhe He, Huangxuan Wu, Junhao Shi, Kexin Huang, Zhaoye Fei, Jingjing Gong, Zuxuan Wu, Yu-Gang Jiang, See-Kiong Ng, Tat-Seng Chua, Xipeng Qiu arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in Vision-Language-Action (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rely on explicit instructions, whereas in real-world interactions, humans rarely issue instructions directly. Effective collaboration requires robots to infer user intentions proactively. In this work, we introduce cross-modal contextual instructions, a new setting where intent is derived from spoken dialogue, environmental sounds, and visual cues rather than explicit commands. To address this new setting, we present RoboOmni, a Perceiver-Thinker-Talker-Executor framework based on end-to-end omni-modal LLMs that unifies intention recognition, interaction confirmation, and action execution. RoboOmni fuses auditory and visual signals spatiotemporally for robust intention recognition, while supporting direct speech interaction. To address the absence of training data for proactive intention recognition in robotic manipulation, we build OmniAction, comprising 140k episodes, 5k+ speakers, 2.4k event sounds, 640 backgrounds, and six contextual instruction types. Experiments in simulation and real-world settings show that RoboOmni surpasses text- and ASR-based baselines in success rate, inference speed, intention recognition, and proactive assistance.

📄 PDF Abstract BibTeX arXiv:2510.23763

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation

2025-11-03 · Heyu Guo, Shanmu Wang, Ruichun Ma, Shiqi Jiang 외 arxiv

Vision-language-action (VLA) models have shown strong generalization for robotic action prediction through large-scale vision-language pretraining. However, most existing models rely solely on RGB cameras, limiting their…

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

2026-05-18 · Ruixiang Zhao, Jie Yang, Zijie Xin, Tianyi Wang 외 arxiv

Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-modal large language models. Existing benchm…

ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding

2026-01-15 · Xueyun Tian, Wei Li, Bingbing Xu, Heng Dong 외 arxiv

Recent Omni-multimodal Large Language Models show promise in unified audio, vision, and text modeling. However, streaming audio-video understanding remains challenging, as existing approaches suffer from disjointed capab…

OmniTact: A Multi-Directional High Resolution Touch Sensor

2020-03-16 · Akhil Padmanabha, Frederik Ebert, Stephen Tian, Roberto Calandra 외

Incorporating touch as a sensing modality for robots can enable finer and more robust manipulation skills. Existing tactile sensors are either flat, have small sensitive fields or only provide low-resolution signals. In …

State EstimationVocal Bursts Intensity Prediction

OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts

2025-03-29 · CVPR 2025 1 · Yuxuan Wang, Yueqian Wang, Bo Chen, Tong Wu 외

The rapid advancement of multi-modal language models (MLLMs) like GPT-4o has propelled the development of Omni language models, designed to process and proactively respond to continuous streams of multi-modal data. Despi…

Streaming video understandingVideo Understanding