paper-with-me

홈 › Papers

Detecting Code-Switching between Turkish-English Language Pair

2018-11-01 · WS 2018 11 · Zeynep Yirmibe{\c{s}}o{\u{g}}lu, G{\"u}l{\c{s}}en Eryi{\u{g}}it

Code-switching (usage of different languages within a single conversation context in an alternative manner) is a highly increasing phenomenon in social media and colloquial usage which poses different challenges for natural language processing. This paper introduces the first study for the detection of Turkish-English code-switching and also a small test data collected from social media in order to smooth the way for further studies. The proposed system using character level n-grams and conditional random fields (CRFs) obtains 95.6{\%} micro-averaged F1-score on the introduced test data set.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information Retrieval

Similar Papers 제목 키워드 기반

A Turkish-German Code-Switching Corpus

2016-05-01 · LREC 2016 5 · {\"O}zlem {\c{C}}etino{\u{g}}lu

Bilingual communities often alternate between languages both in spoken and written communication. One such community, Germany residents of Turkish origin produce Turkish-German code-switching, by heavily mixing two langu…

Language IdentificationSentence

Automatic Detection of Code-switching Style from Acoustics

2018-07-01 · WS 2018 7 · Rallab, SaiKrishna i, Sunayana Sitaram, Alan W. black

Multilingual speakers switch between languages in an non-trivial fashion displaying inter sentential, intra sentential, and congruent lexicalization based transitions. While monolingual ASR systems may be capable of reco…

Automatic Speech Recognition (ASR)Language Identificationspeech-recognitionSpeech Recognition

TuGeBiC: A Turkish German Bilingual Code-Switching Corpus

2022-05-02 · Jeanine Treffers-Daller and, Ozlem Çetinoğlu

In this paper we describe the process of collection, transcription, and annotation of recordings of spontaneous speech samples from Turkish-German bilinguals, and the compilation of a corpus called TuGeBiC. Participants …

Language Identification

Part of Speech Annotation of a Turkish-German Code-Switching Corpus

2016-08-01 · WS 2016 8 · {\"O}zlem {\c{C}}etino{\u{g}}lu, {\c{C}}a{\u{g}}r{\i} {\c{C}}{\"o}ltekin
Language Identification

A Code-Switching Corpus of Turkish-German Conversations

2017-04-01 · WS 2017 4 · {\"O}zlem {\c{C}}etino{\u{g}}lu

We present a code-switching corpus of Turkish-German that is collected by recording conversations of bilinguals. The recordings are then transcribed in two layers following speech and orthography conventions, and annotat…

Automatic Speech Recognition (ASR)Language IdentificationLanguage ModellingPart-Of-Speech Tagging+3