paper-with-me

Papers

Enhancing Speech Instruction Understanding and Disambiguation in Robotics via Speech Prosody

2025-06-01 · David Sasu, Kweku Andoh Yamoah, Benedict Quartey, Natalie Schluter

Enabling robots to accurately interpret and execute spoken language instructions is essential for effective human-robot collaboration. Traditional methods rely on speech recognition to transcribe speech into text, often discarding crucial prosodic cues needed for disambiguating intent. We propose a novel approach that directly leverages speech prosody to infer and resolve instruction intent. Predicted intents are integrated into large language models via in-context learning to disambiguate and select appropriate task plans. Additionally, we present the first ambiguous speech dataset for robotics, designed to advance research in speech disambiguation. Our method achieves 95.79% accuracy in detecting referent intents within an utterance and determines the intended task plan of ambiguous instructions with 71.96% accuracy, demonstrating its potential to significantly improve human-robot communication.

📄 PDF Abstract BibTeX arXiv:2506.02057

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality

2026-02-03 · Yoonsang Kim, Divyansh Pradhan, Devshree Jadeja, Arie Kaufman arxiv

We introduce Speech-to-Spatial, a referent disambiguation framework that converts verbal remote-assistance instructions into spatially grounded AR guidance. Unlike prior systems that rely on additional cues (e.g., gestur…

BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs

2025-09-30 · Yue Wang, Ruotian Ma, Xingyu Chen, Zhengliang Shi 외 arxiv

The rise of Large Language Models (LLMs) is reshaping multimodel models, with speech synthesis being a prominent application. However, existing approaches often underutilize the linguistic intelligence of these models, t…

Speech Synthesis

Integrating Disambiguation and User Preferences into Large Language Models for Robot Motion Planning

2024-04-22 · Mohammed Abugurain, Shinkyu Park

This paper presents a framework that can interpret humans' navigation commands containing temporal elements and directly translate their natural language instructions into robot motion planning. Central to our framework …

Motion Planning

DEBATE: A Dataset for Disentangling Textual Ambiguity in Mandarin Through Speech

2025-06-09 · Haotian Guo, Jing Han, Yongfeng Tu, Shihao Gao 외

Despite extensive research on textual and visual disambiguation, disambiguation through speech (DTS) remains underexplored. This is largely due to the lack of high-quality datasets that pair spoken sentences with richly …

UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions

2023-10-04 · Siddhant Arora, Hayato Futami, Jee-weon Jung, Yifan Peng 외

Recent studies leverage large language models with multi-tasking capabilities, using natural language prompts to guide the model's behavior and surpassing performance of task-specific models. Motivated by this, we ask: c…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Multi-Task Learningspeech-recognition+2