paper-with-me

홈 › Papers

A Code-Switching Corpus of Turkish-German Conversations

2017-04-01 · WS 2017 4 · {\"O}zlem {\c{C}}etino{\u{g}}lu

We present a code-switching corpus of Turkish-German that is collected by recording conversations of bilinguals. The recordings are then transcribed in two layers following speech and orthography conventions, and annotated with sentence boundaries and intersentential, intrasentential, and intra-word switch points. The total amount of data is 5 hours of speech which corresponds to 3614 sentences. The corpus aims at serving as a resource for speech or text analysis, as well as a collection for linguistic inquiries.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech Recognition (ASR)Language IdentificationLanguage ModellingPart-Of-Speech TaggingSentenceSentiment AnalysisSpeech Recognition

Similar Papers 제목 키워드 기반

TuGeBiC: A Turkish German Bilingual Code-Switching Corpus

2022-05-02 · Jeanine Treffers-Daller and, Ozlem Çetinoğlu

In this paper we describe the process of collection, transcription, and annotation of recordings of spontaneous speech samples from Turkish-German bilinguals, and the compilation of a corpus called TuGeBiC. Participants …

Language Identification

A Turkish-German Code-Switching Corpus

2016-05-01 · LREC 2016 5 · {\"O}zlem {\c{C}}etino{\u{g}}lu

Bilingual communities often alternate between languages both in spoken and written communication. One such community, Germany residents of Turkish origin produce Turkish-German code-switching, by heavily mixing two langu…

Language IdentificationSentence

Part of Speech Annotation of a Turkish-German Code-Switching Corpus

2016-08-01 · WS 2016 8 · {\"O}zlem {\c{C}}etino{\u{g}}lu, {\c{C}}a{\u{g}}r{\i} {\c{C}}{\"o}ltekin
Language Identification

Subword-Level Language Identification for Intra-Word Code-Switching

2019-04-03 · NAACL 2019 6 · Manuel Mager, Özlem Çetinoğlu, Katharina Kann

Language identification for code-switching (CS), the phenomenon of alternating between two or more languages in conversations, has traditionally been approached under the assumption of a single language per token. Howeve…

Language Identification

Anonymising the SAGT Speech Corpus and Treebank

2022-06-01 · LREC 2022 6 · Özlem Çetinoğlu, Antje Schweitzer

Anonymisation, that is identifying and neutralising sensitive references, is a crucial part of dataset creation. In this paper, we describe the anonymisation process of a Turkish-German code-switching corpus, namely SAGT…