paper-with-me

홈 › Papers

Classification of Spontaneous and Scripted Speech for Multilingual Audio

2024-12-16 · Shahar Elisha, Andrew McDowell, Mariano Beguerisse-Díaz, Emmanouil Benetos

Distinguishing scripted from spontaneous speech is an essential tool for better understanding how speech styles influence speech processing research. It can also improve recommendation systems and discovery experiences for media users through better segmentation of large recorded speech catalogues. This paper addresses the challenge of building a classifier that generalises well across different formats and languages. We systematically evaluate models ranging from traditional, handcrafted acoustic and prosodic features to advanced audio transformers, utilising a large, multilingual proprietary podcast dataset for training and validation. We break down the performance of each model across 11 language groups to evaluate cross-lingual biases. Our experimental analysis extends to publicly available datasets to assess the models' generalisability to non-podcast domains. Our results indicate that transformer-based models consistently outperform traditional feature-based techniques, achieving state-of-the-art performance in distinguishing between scripted and spontaneous speech across various languages.

📄 PDF Abstract BibTeX arXiv:2412.11896

Code (0)

등록된 구현이 없습니다.

Tasks

Recommendation Systems

Similar Papers 제목 키워드 기반

A Novel Scheme to classify Read and Spontaneous Speech

2023-06-13 · Sunil Kumar Kopparapu

The COVID-19 pandemic has led to an increased use of remote telephonic interviews, making it important to distinguish between scripted and spontaneous speech in audio recordings. In this paper, we propose a novel scheme …

AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages

2026-04-09 · Lilian Wanzare, Cynthia Amol, Ezekiel Maina, Nelson Odhiambo 외 arxiv

AfriVoices-KE is a large-scale multilingual speech dataset comprising approximately 3,000 hours of audio across five Kenyan languages: Dholuo, Kikuyu, Kalenjin, Maasai, and Somali. The dataset includes 750 hours of scrip…

Speech Recognition

Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition

2026-09-09 · Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin, Naomi Harte arxiv

Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversation, which involves overlapping speech, …

Audio-Visual Speech Recognition

Generating coherent spontaneous speech and gesture from text

2021-01-14 · Simon Alexanderson, Éva Székely, Gustav Eje Henter, Taras Kucherenko 외

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating…

Gesture GenerationMotion GenerationSpeech Synthesistext-to-speech+1

Spontaneous Informal Speech Dataset for Punctuation Restoration

2024-09-17 · Xing Yi Liu, Homayoon Beigi

Presently, punctuation restoration models are evaluated almost solely on well-structured, scripted corpora. On the other hand, real-world ASR systems and post-processing pipelines typically apply towards spontaneous spee…

Punctuation Restoration