paper-with-me

Papers

PAREDA: A Multi-Accent Speech Dataset of Natural Language Processing Research Discussions

2026-05-18 · Sicheng Jin, Dipankar Srirag, Aditya Joshi arxiv

While modern Automatic Speech Recognition (ASR) systems achieve high accuracy on benchmark corpora, their performance often degrades when there is real-world variability. This work focuses on variability arising due to accented, spontaneous, and domain-specific speech. In particular, we introduce PAper REading DAtaset (PAREDA), a first-of-its-kind multi-accent speech dataset consisting of discussions on academic Natural Language Processing (NLP) papers between speakers with Australian, Indian-English, and Chinese English accents. Each session elicits a spontaneous monologue (a summary of a paper's abstract) and a non-monologue (a question-and-answer session between participants), resulting in a corpus rich with technical jargon and conversational phenomena. We evaluate the performance of SOTA ASR models on PAREDA, analysing the impact of accent mixing and increased speech rate. Our results show that, in the zero-shot setting, models perform worse, confirming the dataset's challenging nature. However, fine-tuning on PAREDA significantly reduces the Word Error Rate (WER), demonstrating that our dataset captures linguistic characteristics often missing from existing corpora. PAREDA serves as a valuable new resource for building and evaluating more robust and inclusive ASR systems for specialised, real-world applications.

📄 PDF Abstract BibTeX arXiv:2605.17860

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors

2026-06-18 · Masaya Kawamura, Yuma Shirahata, Kentaro Mitsui, Reo Shimizu arxiv

Existing mean opinion score (MOS) prediction models typically predict utterance-level naturalness MOS and can be insensitive to localized pitch-accent errors. We propose Pitch-Accent-focused Speech Quality Assessment (PA…

DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech

2024-10-17 · Jan Melechovsky, Ambuj Mehrish, Berrak Sisman, Dorien Herremans

Recent advancements in Text-to-Speech (TTS) systems have enabled the generation of natural and expressive speech from textual input. Accented TTS aims to enhance user experience by making the synthesized speech more rela…

DisentanglementQuantizationSpeech Synthesistext-to-speech+1

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations

2026-06-24 · Ram Annamdevula, Ankit Tatawat, Ashishkumar P. Gudmalwar, Nirmesh J. Shah 외 arxiv

Accent conversion and controllability remain fundamental challenges in cross-lingual text-to-speech (TTS), particularly for low-resource and phonetically diverse Indic languages. While recent large language model (LLM)-b…

Explicit Intensity Control for Accented Text-to-speech

2022-10-27 · Rui Liu, Haolin Zuo, De Hu, Guanglai Gao 외

Accented text-to-speech (TTS) synthesis seeks to generate speech with an accent (L2) as a variant of the standard version (L1). How to control the intensity of accent in the process of TTS is a very interesting research …

speech-recognitionSpeech Recognitiontext-to-speechText to Speech

Analysis of French Phonetic Idiosyncrasies for Accent Recognition

2021-10-18 · Pierre Berjon, Avishek Nag, Soumyabrata Dev

Speech recognition systems have made tremendous progress since the last few decades. They have developed significantly in identifying the speech of the speaker. However, there is a scope of improvement in speech recognit…

Multi-class Classificationspeech-recognitionSpeech Recognition