Sparse Transcription
The transcription bottleneck is often cited as a major obstacle for efforts to document the world’s endangered languages and supply them with language technologies. One solution is to extend methods from automatic speech recognition and machine translation, and recruit linguists to provide narrow phonetic transcriptions and sentence-aligned translations. However, I believe that these approaches are not a good fit with the available data and skills, or with long-established practices that are essentially word-based. In seeking a more effective approach, I consider a century of transcription practice and a wide range of computational approaches, before proposing a computational model based on spoken term detection that I call “sparse transcription.” This represents a shift away from current assumptions that we transcribe phones, transcribe fully, and transcribe first. Instead, sparse transcription combines the older practice of word-level transcription with interpretive, iterative, and interactive processes that are amenable to wider participation and that open the way to new methods for processing oral languages.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSentencespeech-recognitionSpeech RecognitionTranslationSimilar Papers 제목 키워드 기반
Learning Sparse Analytic Filters for Piano Transcription
In recent years, filterbank learning has become an increasingly popular strategy for various audio-related machine learning tasks. This is partly due to its ability to discover task-specific audio characteristics which c…
Information RetrievalMusic Information RetrievalRetrievalMusical Rhythm Transcription Based on Bayesian Piece-Specific Score Models Capturing Repetitions
Most work on musical score models (a.k.a. musical language models) for music transcription has focused on describing the local sequential dependence of notes in musical scores and failed to capture their global repetitiv…
Computational EfficiencyLanguage ModellingMusic TranscriptionRhythmClosed-Loop Transcription via Convolutional Sparse Coding
Autoencoding has achieved great empirical success as a framework for learning generative models for natural images. Autoencoders often use generic deep networks as the encoder or decoder, which are difficult to interpret…
Rolling Shutter CorrectionDJ Mix Transcription with Multi-Pass Non-Negative Matrix Factorization
DJ mix transcription is a crucial step towards DJ mix reverse engineering, which estimates the set of parameters and audio effects applied to a set of existing tracks to produce a performative DJ mix. We introduce a new …
Dynamic Time WarpingSpoken Term Detection Methods for Sparse Transcription in Very Low-resource Settings
We investigate the efficiency of two very different spoken term detection approaches for transcription when the available data is insufficient to train a robust ASR system. This work is grounded in very low-resource lang…
Dynamic Time WarpingPhoneme Recognition