paper-with-me

홈 › Papers

Universal Automatic Phonetic Transcription into the International Phonetic Alphabet

2023-08-07 · Chihiro Taguchi, Yusuke Sakai, Parisa Haghani, David Chiang

This paper presents a state-of-the-art model for transcribing speech in any language into the International Phonetic Alphabet (IPA). Transcription of spoken languages into IPA is an essential yet time-consuming process in language documentation, and even partially automating this process has the potential to drastically speed up the documentation of endangered languages. Like the previous best speech-to-IPA model (Wav2Vec2Phoneme), our model is based on wav2vec 2.0 and is fine-tuned to predict IPA from audio input. We use training data from seven languages from CommonVoice 11.0, transcribed into IPA semi-automatically. Although this training dataset is much smaller than Wav2Vec2Phoneme's, its higher quality lets our model achieve comparable or better results. Furthermore, we show that the quality of our universal speech-to-IPA models is close to that of human annotators.

📄 PDF Abstract BibTeX arXiv:2308.03917

Code (1)

ctaguchi/multipa 공식 구현

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

AlloVera: A Multilingual Allophone Database

2020-04-17 · LREC 2020 5 · David R. Mortensen, Xinjian Li, Patrick Littell, Alexis Michaud 외

We introduce a new resource, AlloVera, which provides mappings from 218 allophones to phonemes for 14 languages. Phonemes are contrastive phonological units, and allophones are their various concrete realizations, which …

speech-recognitionSpeech Recognition

A Universal System for Automatic Text-to-Phonetics Conversion

2019-09-01 · RANLP 2019 9 · Chen Gafni

This paper describes an automatic text-to-phonetics conversion system. The system was constructed to primarily serve as a research tool. It is implemented in a general-purpose linguistic software, which allows it to be i…

Selective Augmentation: Improving Universal Automatic Phonetic Transcription via G2P Bootstrapping

2026-04-29 · Tobias Bystrich, Julia M. Pritzen, Christoph A. Schmidt, Claudia Wich-Reif arxiv

In the field of universal automatic phonetic transcription (APT), clean and diverse training transcriptions are required. However, such high-quality data is limited. We propose the bootstrapping approach Selective Augmen…

IPA Transcription of Bengali Texts

2024-03-29 · Kanij Fatema, Fazle Dawood Haider, Nirzona Ferdousi Turpa, Tanveer Azmal 외

The International Phonetic Alphabet (IPA) serves to systematize phonemes in language, enabling precise textual representation of pronunciation. In Bengali phonology and phonetics, ongoing scholarly deliberations persist …

LAMA-UT: Language Agnostic Multilingual ASR through Orthography Unification and Language-Specific Transliteration

2024-12-19 · Sangmin Lee, Woo-Jin Chung, Hong-Goo Kang

Building a universal multilingual automatic speech recognition (ASR) model that performs equitably across languages has long been a challenge due to its inherent difficulties. To address this task we introduce a Language…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1