paper-with-me

Papers

Augmented Prompt Selection for Evaluation of Spontaneous Speech Synthesis

2020-05-01 · LREC 2020 5 · Eva Szekely, Jens Edlund, Joakim Gustafson

By definition, spontaneous speech is unscripted and created on the fly by the speaker. It is dramatically different from read speech, where the words are authored as text before they are spoken. Spontaneous speech is emergent and transient, whereas text read out loud is pre-planned. For this reason, it is unsuitable to evaluate the usability and appropriateness of spontaneous speech synthesis by having it read out written texts sampled from for example newspapers or books. Instead, we need to use transcriptions of speech as the target - something that is much less readily available. In this paper, we introduce Starmap, a tool allowing developers to select a varied, representative set of utterances from a spoken genre, to be used for evaluation of TTS for a given domain. The selection can be done from any speech recording, without the need for transcription. The tool uses interactive visualisation of prosodic features with t-SNE, along with a tree-based algorithm to guide the user through thousands of utterances and ensure coverage of a variety of prompts. A listening test has shown that with a selection of genre-specific utterances, it is possible to show significant differences across genres between two synthetic voices built from spontaneous speech.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

BEA-Base: A Benchmark for ASR of Spontaneous Hungarian

2022-02-01 · P. Mihajlik, A. Balog, T. E. Gráczi, A. Kohári 외

Hungarian is spoken by 15 million people, still, easily accessible Automatic Speech Recognition (ASR) benchmark datasets - especially for spontaneous speech - have been practically unavailable. In this paper, we introduc…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3

BEA-Base: A Benchmark for ASR of Spontaneous Hungarian

2022-06-01 · LREC 2022 6 · Peter Mihajlik, Andras Balog, Tekla Etelka Graczi, Anna Kohari 외

Hungarian is spoken by 15 million people, still, easily accessible Automatic Speech Recognition (ASR) benchmark datasets – especially for spontaneous speech – have been practically unavailable. In this paper, we introduc…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3

Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS Challenge

2024-12-05 · Fuxiang Tao, Bahman Mirheidari, Madhurananda Pahar, Sophie Young 외

Dementia is associated with various cognitive impairments and typically manifests only after significant progression, making intervention at this stage often ineffective. To address this issue, the Prediction and Recogni…

Impact of Speech Mode in Automatic Pathological Speech Detection

2024-06-14 · Shakeel A. Sheikh, Ina Kodrasi

Automatic pathological speech detection approaches yield promising results in identifying various pathologies. These approaches are typically designed and evaluated for phonetically-controlled speech scenarios, where spe…

Navigate

Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction

2026-05-15 · Si-Belkacem Yamine Ketir, Lenard Paulo Tamayo, Shohei Hisada, Shaowen Peng 외 arxiv

Accurate assessment of cognitive decline from spontaneous speech remains challenging due to limited dataset size and class imbalance. In this work, we propose a large language model (LLM)-driven data augmentation framewo…

Data Augmentation