paper-with-me

홈 › Papers

Acoustically Precise Hesitation Tagging Is Essential for End-to-End Verbatim Transcription Systems

2025-06-04 · Jhen-Ke Lin, Hao-Chien Lu, Chung-Chun Wang, Hong-Yun Lin, Berlin Chen

Verbatim transcription for automatic speaking assessment demands accurate capture of disfluencies, crucial for downstream tasks like error analysis and feedback. However, many ASR systems discard or generalize hesitations, losing important acoustic details. We fine-tune Whisper models on the Speak & Improve 2025 corpus using low-rank adaptation (LoRA), without recourse to external audio training data. We compare three annotation schemes: removing hesitations (Pure), generic tags (Rich), and acoustically precise fillers inferred by Gemini 2.0 Flash from existing audio-transcript pairs (Extra). Our challenge system achieved 6.47% WER (Pure) and 5.81% WER (Extra). Post-challenge experiments reveal that fine-tuning Whisper Large V3 Turbo with the "Extra" scheme yielded a 5.5% WER, an 11.3% relative improvement over the "Pure" scheme (6.2% WER). This demonstrates that explicit, realistic filled-pause labeling significantly enhances ASR accuracy for verbatim L2 speech transcription.

📄 PDF Abstract BibTeX arXiv:2506.04076

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR

2026-06-12 · Henri-Leon Kordt, Theresa Pekarek Rosin, Jae Hee Lee, Stefan Wermter arxiv

Despite advances in large-scale Automatic Speech Recognition (ASR), disfluent speech remains challenging, as state-of-the-art systems are often optimized to omit disfluencies, leading to information loss and hallucinatio…

Continual LearningSpeech Recognition

Towards Interactive Annotation for Hesitation in Conversational Speech

2020-05-01 · LREC 2020 5 · Jane Wottawa, Marie Tahon, Apolline Marin, Nicolas Audibert

Manual annotation of speech corpora is expensive in both human resources and time. Furthermore, recognizing affects in spontaneous, non acted speech presents a challenge for humans and machines. The aim of the present st…

Reliable Part-of-Speech Tagging of Historical Corpora through Set-Valued Prediction

2020-08-04 · Stefan Heid, Marcel Wever, Eyke Hüllermeier

Syntactic annotation of corpora in the form of part-of-speech (POS) tags is a key requirement for both linguistic research and subsequent automated natural language processing (NLP) tasks. This problem is commonly tackle…

Part-Of-Speech TaggingPOSPOS TaggingTAG

HESITA(te) in Portuguese

2014-05-01 · LREC 2014 5 · C, Sara eias, Dirce Celorico, Jorge Proen{\c{c}}a 외

Hesitations, so-called disfluencies, are a characteristic of spontaneous speech, playing a primary role in its structure, reflecting aspects of the language production and the management of inter-communication. In this p…

Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Management+3

Geo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context

2026-03-11 · Yuanbo Hou, Yanru Wu, Qiaoqiao Ren, Shengchen Li 외 arxiv

Environmental sound understanding in computational auditory scene analysis (CASA) is often formulated as an audio-only recognition problem. This formulation leaves a persistent drawback in multi-label audio tagging (AT):…

Audio Tagging