A Turkish-German Code-Switching Corpus
Bilingual communities often alternate between languages both in spoken and written communication. One such community, Germany residents of Turkish origin produce Turkish-German code-switching, by heavily mixing two languages at discourse, sentence, or word level. Code-switching in general, and Turkish-German code-switching in particular, has been studied for a long time from a linguistic perspective. Yet resources to study them from a more computational perspective are limited due to either small size or licence issues. In this work we contribute the solution of this problem with a corpus. We present a Turkish-German code-switching corpus which consists of 1029 tweets, with a majority of intra-sentential switches. We share different type of code-switching we have observed in our collection and describe our processing steps. The first step is data collection and filtering. This is followed by manual tokenisation and normalisation. And finally, we annotate data with word-level language identification information. The resulting corpus is available for research purposes.
Code (0)
등록된 구현이 없습니다.
Tasks
Language IdentificationSentenceSimilar Papers 제목 키워드 기반
A Code-Switching Corpus of Turkish-German Conversations
We present a code-switching corpus of Turkish-German that is collected by recording conversations of bilinguals. The recordings are then transcribed in two layers following speech and orthography conventions, and annotat…
Automatic Speech Recognition (ASR)Language IdentificationLanguage ModellingPart-Of-Speech Tagging+3TuGeBiC: A Turkish German Bilingual Code-Switching Corpus
In this paper we describe the process of collection, transcription, and annotation of recordings of spontaneous speech samples from Turkish-German bilinguals, and the compilation of a corpus called TuGeBiC. Participants …
Language IdentificationPart of Speech Annotation of a Turkish-German Code-Switching Corpus
Anonymising the SAGT Speech Corpus and Treebank
Anonymisation, that is identifying and neutralising sensitive references, is a crucial part of dataset creation. In this paper, we describe the anonymisation process of a Turkish-German code-switching corpus, namely SAGT…
Universal Dependencies Treebank for Tatar: Incorporating Intra-Word Code-Switching Information
This paper introduces a new Universal Dependencies treebank for the Tatar language named NMCTT. A significant feature of the corpus is that it includes code-switching (CS) information at a morpheme level, given the fact …
Language IdentificationPOSTAG