paper-with-me

홈 › Papers

WhisperAlign: Word-Boundary-Aware ASR and WhisperX-Anchored Pyannote Diarization for Long-Form Bengali Speech

2026-03-05 · Aurchi Chowdhury, Rubaiyat -E-Zaman, Sk. Ashrafuzzaman Nafees arxiv

This paper presents our solution for the DL Sprint 4.0, addressing the dual challenges of Bengali Long-Form Speech Recognition (Task 1) and Speaker Diarization (Task 2). Processing long-form, multi-speaker Bengali audio introduces significant hurdles in voice activity detection, overlapping speech, and context preservation. To solve the long-form transcription challenge, we implemented a robust audio chunking strategy utilizing whisper-timestamped, allowing us to feed precise, context-aware segments into our fine-tuned acoustic model for high-accuracy transcription. For the diarization task, we developed an integrated pipeline leveraging pyannote.audio and WhisperX. A key contribution of our approach is the domain-specific fine-tuning of the Pyannote segmentation model on the competition dataset. This adaptation allowed the model to better capture the nuances of Bengali conversational dynamics and accurately resolve complex, overlapping speaker boundaries. Our methodology demonstrates that applying intelligent timestamped chunking to ASR and targeted segmentation fine-tuning to diarization significantly drives down Word Error Rate (WER) and Diarization Error Rate (DER), in low-resource settings.

📄 PDF Abstract BibTeX arXiv:2603.04809

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker DiarizationSpeech RecognitionActivity Detection

Similar Papers 제목 키워드 기반

OxfordVGG Submission to the EGO4D AV Transcription Challenge

2023-07-18 · Jaesung Huh, Max Bain, Andrew Zisserman

This report presents the technical details of our submission on the EGO4D Audio-Visual (AV) Automatic Speech Recognition Challenge 2023 from the OxfordVGG team. We present WhisperX, a system for efficient speech transcri…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Tradition or Innovation: A Comparison of Modern ASR Methods for Forced Alignment

2024-06-27 · Rotem Rousso, Eyal Cohen, Joseph Keshet, Eleanor Chodroff

Forced alignment (FA) plays a key role in speech research through the automatic time alignment of speech signals with corresponding text transcriptions. Despite the move towards end-to-end architectures for speech techno…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Evaluating ASR robustness to spontaneous speech errors: A study of WhisperX using a Speech Error Database

2025-08-18 · John Alderete, Macarious Kin Fung Hui, Aanchan Mohan arxiv

The Simon Fraser University Speech Error Database (SFUSED) is a public data collection developed for linguistic and psycholinguistic research. Here we demonstrate how its design and annotations can be used to test and ev…

Speech Recognition

Boundary-Aware NL2SQL: Integrating Reliability through Hybrid Reward and Data Synthesis

2026-01-15 · Songsong Tian, Kongsheng Zhuo, Zhendong Wang, Rong Shen 외 arxiv

In this paper, we present BAR-SQL (Boundary-Aware Reliable NL2SQL), a unified training framework that embeds reliability and boundary awareness directly into the generation process. We introduce a Seed Mutation data synt…

Reinforcement Learning

Geometry-Aware Infrastructure-Anchored Denoiser for UWB Sensing and Work-Zone Reconstruction

2026-07-05 · Weizhe Tang, Jiaxi Liu, Junwei you, Steven T. Parker 외 arxiv

Accurate work-zone geometry perception is critical for intelligent transportation systems, and ultra-wideband sensing offers a low-cost approach for infrastructure-aided reconstruction. However, outdoor UWB ranging is of…