paper-with-me

홈 › Papers

Nunc profana tractemus. Detecting Code-Switching in a Large Corpus of 16th Century Letters

2022-06-01 · LREC 2022 6 · Martin Volk, Lukas Fischer, Patricia Scheurer, Bernard Silvan Schroffenegger, Raphael Schwitter, Phillip Ströbel, Benjamin Suter

This paper is based on a collection of 16th century letters from and to the Zurich reformer Heinrich Bullinger. Around 12,000 letters of this exchange have been preserved, out of which 3100 have been professionally edited, and another 5500 are available as provisional transcriptions. We have investigated code-switching in these 8600 letters, first on the sentence-level and then on the word-level. In this paper we give an overview of the corpus and its language mix (mostly Early New High German and Latin, but also French, Greek, Italian and Hebrew). We report on our experiences with a popular language identifier and present our results when training an alternative identifier on a very small training corpus of only 150 sentences per language. We use the automatically labeled sentences in order to bootstrap a word-based language classifier which works with high accuracy. Our research around the corpus building and annotation involves automatic handwritten text recognition, text normalisation for ENH German, and machine translation from medieval Latin into modern German.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Handwritten Text RecognitionMachine TranslationSentence

Similar Papers 제목 키워드 기반

Pronunciation Generation for Foreign Language Words in Intra-Sentential Code-Switching Speech Recognition

2022-10-26 · Wei Wang, Chao Zhang, Xiaopei Wu

Code-Switching refers to the phenomenon of switching languages within a sentence or discourse. However, limited code-switching , different language phoneme-sets and high rebuilding costs throw a challenge to make the spe…

Sentencespeech-recognitionSpeech Recognition

Exploring Retraining-Free Speech Recognition for Intra-sentential Code-Switching

2021-08-27 · Zhen Huang, Xiaodan Zhuang, Daben Liu, Xiaoqiang Xiao 외

In this paper, we present our initial efforts for building a code-switching (CS) speech recognition system leveraging existing acoustic models (AMs) and language models (LMs), i.e., no training required, and specifically…

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Leveraging Language ID to Calculate Intermediate CTC Loss for Enhanced Code-Switching Speech Recognition

2023-12-15 · Tzu-Ting Yang, Hsin-Wei Wang, Berlin Chen

In recent years, end-to-end speech recognition has emerged as a technology that integrates the acoustic, pronunciation dictionary, and language model components of the traditional Automatic Speech Recognition model. It i…

Automatic Speech RecognitionLanguage IdentificationLanguage ModelingLanguage Modelling+2

Hindi-English Code-Switching Speech Corpus

2018-09-24 · Sreeram Ganji, Dhawan Kunal, Sinha Rohit

Code-switching refers to the usage of two languages within a sentence or discourse. It is a global phenomenon among multilingual communities and has emerged as an independent area of research. With the increasing demand …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationLanguage Modeling+5

Computer-assisted Pronunciation Training -- Speech synthesis is almost all you need

2022-07-02 · Daniel Korzekwa, Jaime Lorenzo-Trueba, Thomas Drugman, Bozena Kostek

The research community has long studied computer-assisted pronunciation training (CAPT) methods in non-native speech. Researchers focused on studying various model architectures, such as Bayesian networks and deep learni…

AllSpeech Synthesistext-to-speechText to Speech