paper-with-me

홈 › Papers

Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification

2026-08-01 · Bogdan Savelyev arxiv

Off-the-shelf LID and letter heuristics over-label Kazakh-Russian social text as mixed: Russian loanwords inside Kazakh look like code-switching under a shared Cyrillic script. We release a document-level gold LID set whose guideline keeps integrated borrowings as Kazakh and reserves mixed for clause-level switches, plus a mixed-only sentiment pool used after LID in a filter-first cascade. On a shared LID test, FastText, Lingua, raw and windowed HeLI, character-trigram NB, and XLM-R range from weak to strong performance. The gap shows the bottleneck is the loanword-vs-switch annotation boundary, not model class alone.

📄 PDF Abstract BibTeX arXiv:2608.00581

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Large Vocabulary Kazakh-Russian Sign Language Dataset: KRSL-OnlineSchool

2022-06-01 · SignLang (LREC) 2022 6 · Medet Mukushev, Aigerim Kydyrbekova, Vadim Kimmelman, Anara Sandygulova

This paper presents a new dataset for Kazakh-Russian Sign Language (KRSL) created for the purposes of Sign Language Processing. In 2020, Kazakhstan’s schools were quickly switched to online mode due to the COVID-19 pande…

Sign Language TranslationSpeech-to-Text

Low-resource Machine Translation for Code-switched Kazakh-Russian Language Pair

2025-03-25 · Maksim Borisov, Zhanibek Kozhirbayev, Valentin Malykh

Machine translation for low resource language pairs is a challenging task. This task could become extremely difficult once a speaker uses code switching. We propose a method to build a machine translation model for code-…

Machine TranslationTranslation

100,000+ Movie Reviews from Kazakhstan: Russian, Kazakh, and Code-Switched Texts

2026-05-09 · Rustem Yeshpanov arxiv

We present a new publicly available corpus of 100,502 movie reviews from Kazakhstan collected from kino.kz, spanning 2001-2025 and covering 4,943 unique titles. The dataset is multilingual, consisting mainly of Russian r…

KazNERD: Kazakh Named Entity Recognition Dataset

2021-11-26 · LREC 2022 6 · Rustem Yeshpanov, Yerbolat Khassanov, Huseyin Atakan Varol

We present the development of a dataset for Kazakh named entity recognition. The dataset was built as there is a clear need for publicly available annotated corpora in Kazakh, as well as annotation guidelines containing …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Acquisition of Translation Lexicons for Historically Unwritten Languages via Bridging Loanwords

2017-06-06 · WS 2017 8 · Michael Bloodgood, Benjamin Strauss

With the advent of informal electronic communications such as social media, colloquial languages that were historically unwritten are being written for the first time in heavily code-switched environments. We present a m…

Machine TranslationTranslation