paper-with-me

Papers

SIFToM: Robust Spoken Instruction Following through Theory of Mind

2024-09-17 · Lance Ying, Jason Xinyu Liu, Shivam Aarya, Yizirui Fang, Stefanie Tellex, Joshua B. Tenenbaum, Tianmin Shu

Spoken language instructions are ubiquitous in agent collaboration. However, in human-robot collaboration, recognition accuracy for human speech is often influenced by various speech and environmental factors, such as background noise, the speaker's accents, and mispronunciation. When faced with noisy or unfamiliar auditory inputs, humans use context and prior knowledge to disambiguate the stimulus and take pragmatic actions, a process referred to as top-down processing in cognitive science. We present a cognitively inspired model, Speech Instruction Following through Theory of Mind (SIFToM), to enable robots to pragmatically follow human instructions under diverse speech conditions by inferring the human's goal and joint plan as prior for speech perception and understanding. We test SIFToM in simulated home experiments (VirtualHome 2). Results show that the SIFToM model outperforms state-of-the-art speech and language models, approaching human-level accuracy on challenging speech instruction following tasks. We then demonstrate its ability at the task planning level on a mobile manipulator for breakfast preparation tasks.

📄 PDF Abstract BibTeX arXiv:2409.10849

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingTask Planning

Similar Papers 제목 키워드 기반

Situated Instruction Following

2024-07-15 · So Yeon Min, Xavi Puig, Devendra Singh Chaplot, Tsung-Yen Yang 외

Language is never spoken in a vacuum. It is expressed, comprehended, and contextualized within the holistic backdrop of the speaker's history, actions, and environment. Since humans are used to communicating efficiently …

Instruction Following

VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions

2025-09-09 · Jun Zhan, Mingyang Han, Yuxuan Xie, Chen Wang 외 arxiv

Spoken language models (SLMs) have emerged as a unified paradigm for speech understanding and generation, enabling natural human machine interaction. However, while most progress has focused on semantic accuracy and inst…

Instruction Following

Multimodal Speech Recognition for Language-Guided Embodied Agents

2023-02-27 · Allen Chang, Xiaoyuan Zhu, Aarav Monga, Seoho Ahn 외

Benchmarks for language-guided embodied agents typically assume text-based instructions, but deployed agents will encounter spoken instructions. While Automatic Speech Recognition (ASR) models can bridge the input gap, e…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

TiCo: Time-Controllable Spoken Dialogue Model

2026-03-23 · Kai-Wei Chang, Wei-Chih Chen, En-Pei Hu, Hung-yi Lee 외 arxiv

We introduce TiCo, a time-controllable spoken dialogue model (SDM) that follows time-constrained instructions (e.g., "Please generate a response lasting about 15 seconds") and generates spoken responses with controllable…

Reinforcement LearningInstruction Following

DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding

2024-06-13 · Suwon Shon, Kwangyoun Kim, Yi-Te Hsu, Prashant Sridhar 외

The integration of pre-trained text-based large language models (LLM) with speech input has enabled instruction-following capabilities for diverse speech tasks. This integration requires the use of a speech encoder, a sp…

Instruction FollowingLanguage ModelingLanguage ModellingLarge Language Model+2