paper-with-me

Papers

Automatic Long Audio Alignment and Confidence Scoring for Conversational Arabic Speech

2014-05-01 · LREC 2014 5 · Mohamed Elmahdy, Mark Hasegawa-Johnson, Eiman Mustafawi

In this paper, a framework for long audio alignment for conversational Arabic speech is proposed. Accurate alignments help in many speech processing tasks such as audio indexing, speech recognizer acoustic model (AM) training, audio summarizing and retrieving, etc. We have collected more than 1,400 hours of conversational Arabic besides the corresponding human generated non-aligned transcriptions. Automatic audio segmentation is performed using a split and merge approach. A biased language model (LM) is trained using the corresponding text after a pre-processing stage. Because of the dominance of non-standard Arabic in conversational speech, a graphemic pronunciation model (PM) is utilized. The proposed alignment approach is performed in two passes. Firstly, a generic standard Arabic AM is used along with the biased LM and the graphemic PM in a fast speech recognition pass. In a second pass, a more restricted LM is generated for each audio segment, and unsupervised acoustic model adaptation is applied. The recognizer output is aligned with the processed transcriptions using Levenshtein algorithm. The proposed approach resulted in an initial alignment accuracy of 97.8-99.0{\%} depending on the amount of disfluencies. A confidence scoring metric is proposed to accept/reject aligner output. Using confidence scores, it was possible to reject the majority of mis-aligned segments resulting in alignment accuracy of 99.0-99.8{\%} depending on the speech domain and the amount of disfluencies.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Phoneme Similarity Matrices to Improve Long Audio Alignment for Automatic Subtitling

2014-05-01 · LREC 2014 5 · Pablo Ruiz, Aitor {\'A}lvarez, Haritz Arzelus

Long audio alignment systems for Spanish and English are presented, within an automatic subtitling application. Language-specific phone decoders automatically recognize audio contents at phoneme level. At the same time, …

DecoderLanguage Modelling

Cross-Attention with Confidence Weighting for Multi-Channel Audio Alignment

2025-09-21 · Ragib Amin Nihal, Benjamin Yen, Takeshi Ashizawa, Kazuhiro Nakadai arxiv

Multi-channel audio alignment is a key requirement in bioacoustic monitoring, spatial audio systems, and acoustic localization. However, existing methods often struggle to address nonlinear clock drift and lack mechanism…

Binary Classification

Iterative pseudo-forced alignment by acoustic CTC loss for self-supervised ASR domain adaptation

2022-10-27 · Fernando López, Jordi Luque

High-quality data labeling from specific domains is costly and human time-consuming. In this work, we propose a self-supervised domain adaptation method, based upon an iterative pseudo-forced alignment algorithm. The pro…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain Adaptationspeech-recognition+1

A Study of Annotation and Alignment Accuracy for Performance Comparison in Complex Orchestral Music

2019-10-16 · Thassilo Gadermaier, Gerhard Widmer

Quantitative analysis of commonalities and differences between recorded music performances is an increasingly common task in computational musicology. A typical scenario involves manual annotation of different recordings…

Confidence-Aware Automated Assessment of Student-Drawn Scientific Models

2026-06-18 · Luyang Fang, Yingchuan Zhang, Jongchan Park, Zhaoji Wang 외 arxiv

Student-generated drawings are widely used in science education to assess learners' conceptual understanding in modeling-based tasks aligned with the Next Generation Science Standards (NGSS). However, scoring such drawin…