LDC Forced Aligner
This paper describes the LDC forced aligner which was designed to align audio and transcripts. Unlike existing forced aligners, LDC forced aligner can align partially transcribed audio files, and also audio files with large chunks of non-speech segments, such as noise, music, silence etc, by inserting optional wildcard phoneme sequences between sentence or paragraph boundaries. Based on the HTK tool kit, LDC forced aligner can align audio and transcript on sentence or word level. This paper also reports its usage on English and Mandarin Chinese data.
Code (0)
등록된 구현이 없습니다.
Tasks
SentenceSpeech RecognitionSpeech SynthesisText-To-Speech SynthesisSimilar Papers 제목 키워드 기반
Grapheme-Based Cross-Language Forced Alignment: Results with Uralic Languages
Forced alignment is an effective process to speed up linguistic research. However, most forced aligners are language-dependent, and under-resourced languages rarely have enough resources to train an acoustic model for an…
The Mason-Alberta Phonetic Segmenter: A forced alignment system based on deep neural networks and interpolation
Forced alignment systems automatically determine boundaries between segments in speech data, given an orthographic transcription. These tools are commonplace in phonetics to facilitate the use of speech data that would b…
Greek Forced Alignment: Assessing the Accuracy of the Montreal Forced Aligner
Forced alignment has allowed for the rapid creation and annotation of corpora. In this study we examine the Montreal Foreced Aligner and its accuracy of aligning Greek data. Using a conversational Greek corpus we train a…
Montreal Forced Aligner and the state of speech-to-text alignment in 2026
The Montreal Forced Aligner (MFA) was released in 2016 and has since become the most widely used tool for forced alignment in research and industry. In the decade since, MFA has undergone substantial development, includi…
Corpus Phonetics Tutorial
Corpus phonetics has become an increasingly popular method of research in linguistic analysis. With advances in speech technology and computational power, large scale processing of speech data has become a viable techniq…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition