paper-with-me

Papers

A Multimodal Data Collection Framework for Dialogue-Driven Assistive Robotics to Clarify Ambiguities: A Wizard-of-Oz Pilot Study

2026-01-23 · Guangping Liu, Nicholas Hawkins, Billy Madden, Tipu Sultan, Flavio Esposito, Madi Babaiasl arxiv

Integrated control of wheelchairs and wheelchair-mounted robotic arms (WMRAs) has strong potential to increase independence for users with severe motor limitations, yet existing interfaces often lack the flexibility needed for intuitive assistive interaction. Although data-driven AI methods show promise, progress is limited by the lack of multimodal datasets that capture natural Human-Robot Interaction (HRI), particularly conversational ambiguity in dialogue-driven control. To address this gap, we propose a multimodal data collection framework that employs a dialogue-based interaction protocol and a two-room Wizard-of-Oz (WoZ) setup to simulate robot autonomy while eliciting natural user behavior. The framework records five synchronized modalities: RGB-D video, conversational audio, inertial measurement unit (IMU) signals, end-effector Cartesian pose, and whole-body joint states across five assistive tasks. Using this framework, we collected a pilot dataset of 53 trials from five participants and validated its quality through motion smoothness analysis and user feedback. The results show that the framework effectively captures diverse ambiguity types and supports natural dialogue-driven interaction, demonstrating its suitability for scaling to a larger dataset for learning, benchmarking, and evaluation of ambiguity-aware assistive control.

📄 PDF Abstract BibTeX arXiv:2601.16870

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MConv: An Environment for Multimodal Conversational Search across Multiple Domains

2021-07-11 · SIGIR 2021 7 · Lizi Liao, Le Hong Long, Zheng Zhang, Minlie Huang 외

Although conversational search has become a hot topic in both dialogue research and IR community, the real breakthrough has been limited by the scale and quality of datasets available. To address this fundamental obstacl…

Conversational RecommendationConversational SearchDialogue State TrackingResponse Generation

A Domain-Specific Language for LLM-Driven Trigger Generation in Multimodal Data Collection

2026-03-13 · Philipp Reis, Philipp Rigoll, Martin Zehetner, Jacqueline Henle 외 arxiv

Data-driven systems depend on task-relevant data, yet data collection pipelines remain passive and indiscriminate. Continuous logging of multimodal sensor streams incurs high storage costs and captures irrelevant data. T…

Code Generation

The REX corpora: A collection of multimodal corpora of referring expressions in collaborative problem solving dialogues

2012-05-01 · LREC 2012 5 · Takenobu Tokunaga, Ryu Iida, Asuka Terai, Naoko Kuriyama

This paper describes a collection of multimodal corpora of referring expressions, the REX corpora. The corpora have two notable features, namely (1) they include time-aligned extra-linguistic information such as particip…

EXMODD: An EXplanatory Multimodal Open-Domain Dialogue dataset

2023-10-17 · Hang Yin, Pinren Lu, Ziang Li, Bin Sun 외

The need for high-quality data has been a key issue hindering the research of dialogue tasks. Recent studies try to build datasets through manual, web crawling, and large pre-trained models. However, man-made data is exp…

Language Modelling

Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling

2024-12-30 · Min Zhang, Zilin Wang, Liyan Chen, KunHong Liu 외

Recent advances in AI-driven storytelling have enhanced video generation and story visualization. However, translating dialogue-centric scripts into coherent storyboards remains a significant challenge due to limited scr…

Retrieval-augmented GenerationStory VisualizationVideo Generation