paper-with-me

홈 › Papers

Speech Recovery for Real-World Self-powered Intermittent Devices

2021-06-09 · Yu-Chen Lin, Tsun-An Hsieh, Kuo-Hsuan Hung, Cheng Yu, Harinath Garudadri, Yu Tsao, Tei-Wei Kuo

The incompleteness of speech inputs severely degrades the performance of all the related speech signal processing applications. Although many researches have been proposed to address this issue, they controlled the data missing conditions by simulation with self-defined masking lengths or sizes. Besides, the masking definitions are different among all these experimental settings. This paper presents a novel intermittent speech recovery (ISR) system for real-world self-powered intermittent devices. Three contributive stages: interpolation, enhancement, and combination are applied to the ISR system for speech reconstruction. The experimental results show that our recovery system increases speech quality by up to 591.7%, while increasing speech intelligibility by up to 80.5%. Most importantly, the proposed ISR system improves the WER scores by up to 52.6%. The promising results not only confirm the effectiveness of the reconstruction but also encourage the utilization of these battery-free wearable/IoT devices.

📄 PDF Abstract BibTeX arXiv:2106.05229

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Self-Powered LLM Modality Expansion for Large Speech-Text Models

2024-10-04 · Tengfei Yu, Xuebo Liu, Zhiyi Hou, Liang Ding 외

Large language models (LLMs) exhibit remarkable performance across diverse tasks, indicating their potential for expansion into large speech-text models (LSMs) by integrating speech capabilities. Although unified speech-…

Automatic Speech RecognitionInstruction Followingspeech-recognitionSpeech Recognition

What Did I Just Say? Self-Listening for Full-Duplex Speech Models

2026-09-04 · Xuanning Zhou, Junyi Ao, Xiaotong Liu, Tom Ko 외 hf

Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions and backchannels in human conversation. However, text generation, speech synthesis, and audio playback proceed…

Speech SynthesisText Generation

RFM-HRI : A Multimodal Dataset of Medical Robot Failure, User Reaction and Recovery Preferences for Item Retrieval Tasks

2026-03-05 · Yashika Batra, Giuliano Pioldi, Promise Ekpo, Arman Sayatqyzy 외 arxiv

While robots deployed in real-world environments inevitably experience interaction failures, understanding how users respond through verbal and non-verbal behaviors remains under-explored in human-robot interaction (HRI)…

Drones that Think on their Feet: Sudden Landing Decisions with Embodied AI

2025-09-30 · Diego Ortiz Barbosa, Mohit Agrawal, Yash Malegaonkar, Luis Burbano 외 arxiv

Autonomous drones must often respond to sudden events, such as alarms, faults, or unexpected changes in their environment, that require immediate and adaptive decision-making. Traditional approaches rely on safety engine…

Boosting keyword spotting through on-device learnable user speech characteristics

2024-03-12 · Cristian Cioflan, Lukas Cavigelli, Luca Benini

Keyword spotting systems for always-on TinyML-constrained applications require on-site tuning to boost the accuracy of offline trained classifiers when deployed in unseen inference conditions. Adapting to the speech pecu…

Few-Shot LearningKeyword Spotting