paper-with-me

Papers

Prominence-aware automatic speech recognition for conversational speech

2025-09-12 · Julian Linke, Barbara Schuppler arxiv

This paper investigates prominence-aware automatic speech recognition (ASR) by combining prominence detection and speech recognition for conversational Austrian German. First, prominence detectors were developed by fine-tuning wav2vec2 models to classify word-level prominence. The detector was then used to automatically annotate prosodic prominence in a large corpus. Based on those annotations, we trained novel prominence-aware ASR systems that simultaneously transcribe words and their prominence levels. The integration of prominence information did not change performance compared to our baseline ASR system, while reaching a prominence detection accuracy of 85.53% for utterances where the recognized word sequence was correct. This paper shows that transformer-based models can effectively encode prosodic information and represents a novel contribution to prosody-enhanced ASR, with potential applications for linguistic research and prosody-informed dialogue systems.

📄 PDF Abstract BibTeX arXiv:2509.10116

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Crowdsourced and Automatic Speech Prominence Estimation

2023-10-12 · Max Morrison, Pranav Pawar, Nathan Pruyne, Jennifer Cole 외

The prominence of a spoken word is the degree to which an average native listener perceives the word as salient or emphasized relative to its context. Speech prominence estimation is the process of assigning a numeric va…

Emotion Recognitiontext-to-speechText to Speech

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models

2025-06-06 · Yuke Lin, Ming Cheng, Ze Li, Beilong Tang 외

Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialize…

Automatic Speech Recognitionspeaker-diarizationSpeaker Diarizationspeech-recognition+1

Speaker-Aware Simulation Improves Conversational Speech Recognition

2026-02-04 · Máté Gedeon, Péter Mihajlik arxiv

Automatic speech recognition (ASR) for conversational speech remains challenging due to the limited availability of large-scale, well-annotated multi-speaker dialogue data and the complex temporal dynamics of natural int…

Dialogue GenerationSpeech RecognitionData Augmentation

Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion

2019-06-27 · ACL 2019 7 · Suyoun Kim, Siddharth Dalmia, Florian Metze

We present a novel conversational-context aware end-to-end speech recognizer based on a gated neural network that incorporates conversational-context/word/speech embeddings. Unlike conventional speech recognition models,…

SentenceSentence Embeddingsspeech-recognitionSpeech Recognition

Evaluation of Automated Speech Recognition Systems for Conversational Speech: A Linguistic Perspective

2022-11-05 · Hannaneh B. Pasandi, Haniyeh B. Pasandi

Automatic speech recognition (ASR) meets more informal and free-form input data as voice user interfaces and conversational agents such as the voice assistants such as Alexa, Google Home, etc., gain popularity. Conversat…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition