paper-with-me

Papers

Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling

2025-02-05 · Jakob Poncelet, Hugo Van hamme

The recent advancement of speech recognition technology has been driven by large-scale datasets and attention-based architectures, but many challenges still remain, especially for low-resource languages and dialects. This paper explores the integration of weakly supervised transcripts from TV subtitles into automatic speech recognition (ASR) systems, aiming to improve both verbatim transcriptions and automatically generated subtitles. To this end, verbatim data and subtitles are regarded as different domains or languages, due to their distinct characteristics. We propose and compare several end-to-end architectures that are designed to jointly model both modalities with separate or shared encoders and decoders. The proposed methods are able to jointly generate a verbatim transcription and a subtitle. Evaluation on Flemish (Belgian Dutch) demonstrates that a model with cascaded encoders and separate decoders allows to represent the differences between the two data types most efficiently while improving on both domains. Despite differences in domain and linguistic variations, combining verbatim transcripts with subtitle data leads to notable ASR improvements without the need for extensive preprocessing. Additionally, experiments with a large-scale subtitle dataset show the scalability of the proposed approach. The methods not only improve ASR accuracy but also generate subtitles that closely match standard written text, offering several potential applications.

📄 PDF Abstract BibTeX arXiv:2502.03212

Code (1)

nelfproject/NeLF_Transcription_ASR 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

The 2015 Sheffield System for Transcription of Multi-Genre Broadcast Media

2015-12-21 · Oscar Saz, Mortaza Doulaty, Salil Deena, Rosanna Milner 외

We describe the University of Sheffield system for participation in the 2015 Multi-Genre Broadcast (MGB) challenge task of transcribing multi-genre broadcast shows. Transcription was one of four tasks proposed in the MGB…

Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Modelling+2

Refining Transcripts With TV Subtitles by Prompt-Based Weakly Supervised Training of ASR

2025-09-01 · Xinnian Zhao, Hugo Van Hamme arxiv

This study proposes a novel approach to using TV subtitles within a weakly supervised (WS) Automatic Speech Recognition (ASR) framework. Although TV subtitles are readily available, their imprecise alignment with corresp…

Speech Recognition

SBAAM! Eliminating Transcript Dependency in Automatic Subtitling

2024-05-17 · Marco Gaido, Sara Papi, Matteo Negri, Mauro Cettolo 외

Subtitling plays a crucial role in enhancing the accessibility of audiovisual content and encompasses three primary subtasks: translating spoken dialogue, segmenting translations into concise textual units, and estimatin…

Automatic Segmentation of Broadcast News Audio using Self Similarity Matrix

2014-03-27 · Sapna Soni, Ahmed Imran, Sunil Kumar Kopparapu

Generally audio news broadcast on radio is com- posed of music, commercials, news from correspondents and recorded statements in addition to the actual news read by the newsreader. When news transcripts are available, au…

Segmentation

Semi-supervised Learning with Constraints for Person Identification in Multimedia Data

2013-06-01 · CVPR 2013 6 · Martin Bauml, Makarand Tapaswi, Rainer Stiefelhagen

We address the problem of person identification in TV series. We propose a unified learning framework for multiclass classification which incorporates labeled and unlabeled data, and constraints between pairs of features…

Face RecognitionGeneral ClassificationPerson Identificationregression