paper-with-me

홈 › Papers

Imitation Learning for Elder-Facing Speech Synthesis

2026-06-19 · Dongrui Han, Weidong Chen, Jiawen Kang, Mingyu Cui, Helen Meng, Xixin Wu arxiv

Recent advances in text-to-speech (TTS) synthesis have achieved highly natural and expressive speech generation. However, these systems are designed for general adults and overlook older adults' speech comprehension needs due to age-related sensory and cognitive decline. Prior work involves older adults by collecting preference feedback to tune model parameters. However, obtaining sufficient preference data is costly and difficult, as older adults quickly become fatigued during collection. In this paper, we propose a novel imitation learning (IL) framework to learn TTS models from expert demonstrations. We further improve Group Relative Policy Optimization (GRPO) with two-stage on-policy reward learning (OPRL) to mitigate reward hacking under limited supervision from expert demonstration. Experimental results show that GRPO w/ OPRL outperforms GRPO and supervised baselines in objective and subjective metrics. Audio samples are available at https://dongru1.github.io/demo/im-efss

📄 PDF Abstract BibTeX arXiv:2606.21053

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

Elderly-Contextual Data Augmentation via Speech Synthesis for Elderly ASR

2026-04-15 · Minsik Lee, Seoi Hong, Chongmin Lee, Sieun Choi 외 arxiv

Despite recent progress in automatic speech recognition (ASR), elderly ASR (EASR) remains challenging due to limited training data and the distinct acoustic and linguistic characteristics of elderly speech. In this work,…

Speech RecognitionData AugmentationSpeech Synthesis

VOTE400(Voide Of The Elderly 400 Hours): A Speech Dataset to Study Voice Interface for Elderly-Care

2021-01-20 · Minsu Jang, Sangwon Seo, Dohyung Kim, Jaeyeon Lee 외

This paper introduces a large-scale Korean speech dataset, called VOTE400, that can be used for analyzing and recognizing voices of the elderly people. The dataset includes about 300 hours of continuous dialog speech and…

speech-recognitionSpeech Recognition

Improving Speech Recognition for the Elderly: A New Corpus of Elderly Japanese Speech and Investigation of Acoustic Modeling for Speech Recognition

2020-05-01 · LREC 2020 5 · Meiko Fukuda, Hiromitsu Nishizaki, Yurie Iribe, Ryota Nishimura 외

In an aging society like Japan, a highly accurate speech recognition system is needed for use in electronic devices for the elderly, but this level of accuracy cannot be obtained using conventional speech recognition sys…

speech-recognitionSpeech Recognition

A Review of Challenges in Speech-based Conversational AI for Elderly Care

2024-12-10 · Willemijn Klaassen, Bram van Dijk, Marco Spruit

Artificially intelligent systems optimized for speech conversation are appearing at a fast pace. Such models are interesting from a healthcare perspective, as these voice-controlled assistants may support the elderly and…

Speech Corpus Spoken by Young-old, Old-old and Oldest-old Japanese

2016-05-01 · LREC 2016 5 · Yurie Iribe, Norihide Kitaoka, Shuhei Segawa

We have constructed a new speech data corpus, using the utterances of 100 elderly Japanese people, to improve speech recognition accuracy of the speech of older people. Humanoid robots are being developed for use in elde…

speech-recognitionSpeech Recognition