paper-with-me

홈 › Papers

TuGeBiC: A Turkish German Bilingual Code-Switching Corpus

2022-05-02 · Jeanine Treffers-Daller and, Ozlem Çetinoğlu

In this paper we describe the process of collection, transcription, and annotation of recordings of spontaneous speech samples from Turkish-German bilinguals, and the compilation of a corpus called TuGeBiC. Participants in the study were adult Turkish-German bilinguals living in Germany or Turkey at the time of recording in the first half of the 1990s. The data were manually tokenised and normalised, and all proper names (names of participants and places mentioned in the conversations) were replaced with pseudonyms. Token-level automatic language identification was performed, which made it possible to establish the proportions of words from each language. The corpus is roughly balanced between both languages. We also present quantitative information about the number of code-switches, and give examples of different types of code-switching found in the data. The resulting corpus has been made freely available to the research community.

📄 PDF Abstract BibTeX arXiv:2205.00868

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identification

Similar Papers 제목 키워드 기반

A Turkish-German Code-Switching Corpus

2016-05-01 · LREC 2016 5 · {\"O}zlem {\c{C}}etino{\u{g}}lu

Bilingual communities often alternate between languages both in spoken and written communication. One such community, Germany residents of Turkish origin produce Turkish-German code-switching, by heavily mixing two langu…

Language IdentificationSentence

A Code-Switching Corpus of Turkish-German Conversations

2017-04-01 · WS 2017 4 · {\"O}zlem {\c{C}}etino{\u{g}}lu

We present a code-switching corpus of Turkish-German that is collected by recording conversations of bilinguals. The recordings are then transcribed in two layers following speech and orthography conventions, and annotat…

Automatic Speech Recognition (ASR)Language IdentificationLanguage ModellingPart-Of-Speech Tagging+3

Part of Speech Annotation of a Turkish-German Code-Switching Corpus

2016-08-01 · WS 2016 8 · {\"O}zlem {\c{C}}etino{\u{g}}lu, {\c{C}}a{\u{g}}r{\i} {\c{C}}{\"o}ltekin
Language Identification

A Language-aware Approach to Code-switched Morphological Tagging

2021-06-01 · NAACL (CALCS) 2021 6 · Şaziye Betül Özateş, Özlem Çetinoğlu

Morphological tagging of code-switching (CS) data becomes more challenging especially when language pairs composing the CS data have different morphological representations. In this paper, we explore a number of ways of …

Morphological Tagging

Subword-Level Language Identification for Intra-Word Code-Switching

2019-04-03 · NAACL 2019 6 · Manuel Mager, Özlem Çetinoğlu, Katharina Kann

Language identification for code-switching (CS), the phenomenon of alternating between two or more languages in conversations, has traditionally been approached under the assumption of a single language per token. Howeve…

Language Identification