Investigating Transcription Normalization in the Faetar ASR Benchmark
We examine the role of transcription inconsistencies in the Faetar Automatic Speech Recognition benchmark, a challenging low-resource ASR benchmark. With the help of a small, hand-constructed lexicon, we conclude that find that, while inconsistencies do exist in the transcriptions, they are not the main challenge in the task. We also demonstrate that bigram word-based language modelling is of no added benefit, but that constraining decoding to a finite lexicon can be beneficial. The task remains extremely difficult.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech RecognitionLanguage ModellingSimilar Papers 제목 키워드 기반
The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language
We introduce the Faetar Automatic Speech Recognition Benchmark, a benchmark corpus designed to push the limits of current approaches to low-resource speech recognition. Faetar, a Franco-Proven\c{c}al variety spoken prima…
Automatic Speech Recognitionspeech-recognitionSpeech RecognitionTyphoon ASR Real-time: FastConformer-Transducer for Thai Automatic Speech Recognition
Large encoder-decoder models like Whisper achieve strong offline transcription but remain impractical for streaming applications due to high latency. However, due to the accessibility of pre-trained checkpoints, the open…
Speech RecognitionMusical Features for Automatic Music Transcription Evaluation
This technical report gives a detailed, formal description of the features introduced in the paper: Adrien Ycart, Lele Liu, Emmanouil Benetos and Marcus T. Pearce. "Investigating the Perceptual Validity of Evaluation Met…
Information RetrievalMusic Information RetrievalMusic TranscriptionRetrievalBranchShine: Compact Raw-Audio-to-IPA Transcription with a RoPE E-Branchformer Encoder
Speech-to-IPA transcription is useful when the desired output is pronunciation rather than orthographic text, but competitive multilingual systems are often large and evaluation is sensitive to normalization choices. Thi…
TS-Net: OCR Trained to Switch Between Text Transcription Styles
Users of OCR systems, from different institutions and scientific disciplines, prefer and produce different transcription styles. This presents a problem for training of consistent text recognition neural networks on real…
Optical Character Recognition (OCR)