paper-with-me

Papers

No More Mumbles: Enhancing Robot Intelligibility through Speech Adaptation

2024-05-15 · Qiaoqiao Ren, Yuanbo Hou, Dick Botteldooren, Tony Belpaeme

Spoken language interaction is at the heart of interpersonal communication, and people flexibly adapt their speech to different individuals and environments. It is surprising that robots, and by extension other digital devices, are not equipped to adapt their speech and instead rely on fixed speech parameters, which often hinder comprehension by the user. We conducted a speech comprehension study involving 39 participants who were exposed to different environmental and contextual conditions. During the experiment, the robot articulated words using different vocal parameters, and the participants were tasked with both recognising the spoken words and rating their subjective impression of the robot's speech. The experiment's primary outcome shows that spaces with good acoustic quality positively correlate with intelligibility and user experience. However, increasing the distance between the user and the robot exacerbated the user experience, while distracting background sounds significantly reduced speech recognition accuracy and user satisfaction. We next built an adaptive voice for the robot. For this, the robot needs to know how difficult it is for a user to understand spoken language in a particular setting. We present a prediction model that rates how annoying the ambient acoustic environment is and, consequentially, how hard it is to understand someone in this setting. Then, we develop a convolutional neural network model to adapt the robot's speech parameters to different users and spaces, while taking into account the influence of ambient acoustics on intelligibility. Finally, we present an evaluation with 27 users, demonstrating superior intelligibility and user experience with adaptive voice parameters compared to fixed voice.

📄 PDF Abstract BibTeX arXiv:2405.09708

Code (1)

qiaoqiao2323/robot-speech-intelligibility 공식 구현 pytorch

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Exploring Sentence Type Effects on the Lombard Effect and Intelligibility Enhancement: A Comparative Study of Natural and Grid Sentences

2023-09-19 · Hongyang Chen, Yuhong Yang, Zhongyuan Wang, Weiping tu 외

This study explores how sentence types affect the Lombard effect and intelligibility enhancement, focusing on comparisons between natural and grid sentences. Using the Lombard Chinese-TIMIT (LCT) corpus and the Enhanced …

Sentence

Body-conductive acoustic sensors in human-robot communication

2012-05-01 · LREC 2012 5 · Panikos Heracleous, Carlos Ishi, Takahiro Miyashita, Norihiro Hagita

In this study, the use of alternative acoustic sensors in human-robot communication is investigated. In particular, a Non-Audible Murmur (NAM) microphone was applied in teleoperating Geminoid HI-1 robot in noisy environm…

DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance

2024-08-26 · Jinhyeok Yang, Junhyeok Lee, Hyeong-Seok Choi, Seunghun Ji 외

Text-to-Speech (TTS) models have advanced significantly, aiming to accurately replicate human speech's diversity, including unique speaker identities and linguistic nuances. Despite these advancements, achieving an optim…

Diversitytext-to-speechText to Speech

Vocal effort modeling in neural TTS for improving the intelligibility of synthetic speech in noise

2022-03-20 · Tuomo Raitio, Petko Petkov, Jiangchuan Li, Muhammed Shifas 외

We present a neural text-to-speech (TTS) method that models natural vocal effort variation to improve the intelligibility of synthetic speech in the presence of noise. The method consists of first measuring the spectral …

text-to-speechText to Speech

Enhancing Speech Intelligibility in Text-To-Speech Synthesis using Speaking Style Conversion

2020-08-13 · Dipjyoti Paul, Muhammed PV Shifas, Yannis Pantazis, Yannis Stylianou

The increased adoption of digital assistants makes text-to-speech (TTS) synthesis systems an indispensable feature of modern mobile devices. It is hence desirable to build a system capable of generating highly intelligib…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1