paper-with-me

홈 › Papers

Investigating Transcription Normalization in the Faetar ASR Benchmark

2025-08-15 · Leo Peckham, Michael Ong, Naomi Nagy, Ewan Dunbar arxiv

We examine the role of transcription inconsistencies in the Faetar Automatic Speech Recognition benchmark, a challenging low-resource ASR benchmark. With the help of a small, hand-constructed lexicon, we conclude that find that, while inconsistencies do exist in the transcriptions, they are not the main challenge in the task. We also demonstrate that bigram word-based language modelling is of no added benefit, but that constraining decoding to a finite lexicon can be beneficial. The task remains extremely difficult.

📄 PDF Abstract BibTeX arXiv:2508.11771

Code (0)

등록된 구현이 없습니다.

Tasks

Speech RecognitionLanguage Modelling

Similar Papers 제목 키워드 기반

The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language

2024-09-12 · Michael Ong, Sean Robertson, Leo Peckham, Alba Jorquera Jimenez de Aberasturi 외

We introduce the Faetar Automatic Speech Recognition Benchmark, a benchmark corpus designed to push the limits of current approaches to low-resource speech recognition. Faetar, a Franco-Proven\c{c}al variety spoken prima…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Typhoon ASR Real-time: FastConformer-Transducer for Thai Automatic Speech Recognition

2026-01-19 · Warit Sirichotedumrong, Adisai Na-Thalang, Potsawee Manakul, Pittawat Taveekitworachai 외 arxiv

Large encoder-decoder models like Whisper achieve strong offline transcription but remain impractical for streaming applications due to high latency. However, due to the accessibility of pre-trained checkpoints, the open…

Speech Recognition

Musical Features for Automatic Music Transcription Evaluation

2020-04-15 · Adrien Ycart, Lele Liu, Emmanouil Benetos, Marcus T. Pearce

This technical report gives a detailed, formal description of the features introduced in the paper: Adrien Ycart, Lele Liu, Emmanouil Benetos and Marcus T. Pearce. "Investigating the Perceptual Validity of Evaluation Met…

Information RetrievalMusic Information RetrievalMusic TranscriptionRetrieval

BranchShine: Compact Raw-Audio-to-IPA Transcription with a RoPE E-Branchformer Encoder

2026-06-22 · Nikhil Navas, Sergio Chevtchenko, Talisson Damiao, Saeed Afshar arxiv

Speech-to-IPA transcription is useful when the desired output is pronunciation rather than orthographic text, but competitive multilingual systems are often large and evaluation is sensitive to normalization choices. Thi…

TS-Net: OCR Trained to Switch Between Text Transcription Styles

2021-03-09 · Jan Kohút, Michal Hradiš

Users of OCR systems, from different institutions and scientific disciplines, prefer and produce different transcription styles. This presents a problem for training of consistent text recognition neural networks on real…

Optical Character Recognition (OCR)