Detecting de minimis Code-Switching in Historical German Books
Code-switching has long interested linguists, with computational work in particular focusing on speech and social media data (Sitaram et al., 2019). This paper contrasts these informal instances of code-switching to its appearance in more formal registers, by examining the mixture of languages in the Deutsches Textarchiv (DTA), a corpus of 1406 primarily German books from the 17th to 19th centuries. We automatically annotate and manually inspect spans of six embedded languages (Latin, French, English, Italian, Spanish, and Greek) in the corpus. We quantitatively analyze the differences between code-switching patterns in these books and those in more typically studied speech and social media corpora. Furthermore, we address the practical task of predicting code-switching from features of the matrix language alone in the DTA corpus. Such classifiers can help reduce errors when optical character recognition or speech transcription is applied to a large corpus with rare embedded languages.
Code (0)
등록된 구현이 없습니다.
Tasks
Optical Character RecognitionOptical Character Recognition (OCR)Similar Papers 제목 키워드 기반
A Turkish-German Code-Switching Corpus
Bilingual communities often alternate between languages both in spoken and written communication. One such community, Germany residents of Turkish origin produce Turkish-German code-switching, by heavily mixing two langu…
Language IdentificationSentenceNunc profana tractemus. Detecting Code-Switching in a Large Corpus of 16th Century Letters
This paper is based on a collection of 16th century letters from and to the Zurich reformer Heinrich Bullinger. Around 12,000 letters of this exchange have been preserved, out of which 3100 have been professionally edite…
Handwritten Text RecognitionMachine TranslationSentenceMultilingual Semantic Parsing And Code-Switching
Extending semantic parsing systems to new domains and languages is a highly expensive, time-consuming process, so making effective use of existing resources is critical. In this paper, we describe a transfer learning met…
Domain AdaptationSemantic ParsingTransfer LearningWord EmbeddingsPart of Speech Annotation of a Turkish-German Code-Switching Corpus
A Code-Switching Corpus of Turkish-German Conversations
We present a code-switching corpus of Turkish-German that is collected by recording conversations of bilinguals. The recordings are then transcribed in two layers following speech and orthography conventions, and annotat…
Automatic Speech Recognition (ASR)Language IdentificationLanguage ModellingPart-Of-Speech Tagging+3