paper-with-me

홈 › Papers

Overcoming Decoder Inconsistencies in Whisper for Dravidian and Low-Resource Languages

2026-06-08 · Chowdam Venkata Kumar, Kumud Tripathi, Pankaj Wasnik arxiv

Multilingual ASR models such as Whisper perform well on high-resource languages but exhibit substantially higher Word Error Rates (WER) for Dravidian languages compared to Indo-Aryan ones. Through linguistic and dataset analysis, we show that Dravidian languages have longer words, higher vocabulary diversity, and lower repetition, resulting in sparse token distributions and frequent character-level substitution errors. Baseline fine-tuning further reveals decoder imbalance between self-attention (linguistic context) and cross-attention (acoustic cues). Although synthetic token-repetition experiments indicate potential gains, they are impractical. Motivated by these observations, we introduce two decoder-level enhancements: Weighted-Attention, which adaptively balances attention sources, and Self-Conditioning, which reinjects intermediate predictions to improve token consistency. Experiments demonstrate consistent WER reductions for low-resource and agglutinative languages.

📄 PDF Abstract BibTeX arXiv:2606.09535

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Tulu Resource for Machine Translation

2024-03-28 · Manu Narayanan, Noëmi Aepli

We present the first parallel dataset for English-Tulu translation. Tulu, classified within the South Dravidian linguistic family branch, is predominantly spoken by approximately 2.5 million individuals in southwestern I…

Machine TranslationNMTTransfer LearningTranslation

Multistage Fine-tuning Strategies for Automatic Speech Recognition in Low-resource Languages

2024-11-07 · Leena G Pillai, Kavya Manohar, Basil K Raju, Elizabeth Sherly

This paper presents a novel multistage fine-tuning strategy designed to enhance automatic speech recognition (ASR) performance in low-resource languages using OpenAI's Whisper model. In this approach we aim to build ASR …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection

2024-06-14 · Haoyu Wang, Guoqiang Hu, Guodong Lin, Wei-Qiang Zhang 외

As a robust and large-scale multilingual speech recognition model, Whisper has demonstrated impressive results in many low-resource and out-of-distribution scenarios. However, its encoder-decoder structure hinders its ap…

Decoderspeech-recognitionSpeech Recognition

Beyond WER: Probing Whisper's Sub-token Decoder Across Diverse Language Resource Levels

2025-09-29 · Siyu Liang, Nicolas Ballier, Gina-Anne Levow, Richard Wright arxiv

While large multilingual automatic speech recognition (ASR) models achieve remarkable performance, the internal mechanisms of the end-to-end pipeline, particularly concerning fairness and efficacy across languages, remai…

Speech Recognition

Unsupervised Machine Translation On Dravidian Languages

2021-03-29 · EACL (DravidianLangTech) 2021 4 · Sai Koneru, Danni Liu, Jan Niehues

Unsupervised neural machine translation (UNMT) is beneficial especially for low resource languages such as those from the Dravidian family. However, UNMT systems tend to fail in realistic scenarios involving actual low r…

Machine TranslationTranslationUnsupervised Machine Translation