paper-with-me

홈 › Papers

Make It Hard to Hear, Easy to Learn: Long-Form Bengali ASR and Speaker Diarization via Extreme Augmentation and Perfect Alignment

2026-02-26 · Sanjid Hasan, Risalat Labib, A H M Fuad, Bayazid Hasan arxiv

Although Automatic Speech Recognition (ASR) in Bengali has seen significant progress, processing long-duration audio and performing robust speaker diarization remain critical research gaps. To address the severe scarcity of joint ASR and diarization resources for this language, we introduce Lipi-Ghor-882, a comprehensive 882-hour multi-speaker Bengali dataset. In this paper, detailing our submission to the DL Sprint 4.0 competition, we systematically evaluate various architectures and approaches for long-form Bengali speech. For ASR, we demonstrate that raw data scaling is ineffective; instead, targeted fine-tuning utilizing perfectly aligned annotations paired with synthetic acoustic degradation (noise and reverberation) emerges as the singular most effective approach. Conversely, for speaker diarization, we observed that global open-source state-of-the-art models (such as Diarizen) performed surprisingly poorly on this complex dataset. Extensive model retraining yielded negligible improvements; instead, strategic, heuristic post-processing of baseline model outputs proved to be the primary driver for increasing accuracy. Ultimately, this work outlines a highly optimized dual pipeline achieving a $\sim$0.019 Real-Time Factor (RTF), establishing a practical, empirically backed benchmark for low-resource, long-form speech processing.

📄 PDF Abstract BibTeX arXiv:2602.23070

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker DiarizationSpeech Recognition

Similar Papers 제목 키워드 기반

Easy attention: A simple attention mechanism for temporal predictions with transformers

2023-08-24 · Marcial Sanchis-Agudo, Yuning Wang, Roger Arnau, Luca Guastoni 외

To improve the robustness of transformer neural networks used for temporal-dynamics prediction of chaotic systems, we propose a novel attention mechanism called easy attention which we demonstrate in time-series reconstr…

Temporal SequencesTime Series

Identifying Hard Noise in Long-Tailed Sample Distribution

2022-07-27 · Xuanyu Yi, Kaihua Tang, Xian-Sheng Hua, Joo-Hwee Lim 외

Conventional de-noising methods rely on the assumption that all samples are independent and identically distributed, so the resultant classifier, though disturbed by noise, can still easily identify the noises as the out…

Philosophy

Advancing Machine-Generated Text Detection from an Easy to Hard Supervision Perspective

2025-11-02 · Chenwang Wu, Yiu-ming Cheung, Bo Han, Defu Lian arxiv

Existing machine-generated text (MGT) detection methods implicitly assume labels as the "golden standard". However, we reveal boundary ambiguity in MGT detection, implying that traditional training paradigms are inexact.…

Knowledge DistillationText Detection

How can neuromorphic hardware attain brain-like functional capabilities?

2023-10-25 · Wolfgang Maass

Research on neuromorphic computing is driven by the vision that we can emulate brain-like computing capability, learning capability, and energy-efficiency in novel hardware. Unfortunately, this vision has so far been pur…

Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context

2025-03-19 · Junyi Ao, Dekun Chen, Xiaohai Tian, Wenjie Feng 외

Large Language Models (LLMs) have recently shown remarkable ability to process not only text but also multimodal inputs such as speech and audio. However, most existing models primarily focus on analyzing input signals u…

Audio captioningAudio Question AnsweringAudio TaggingQuestion Answering