paper-with-me

홈 › Papers

Ukrainian Texts Classification: Exploration of Cross-lingual Knowledge Transfer Approaches

2024-04-02 · Daryna Dementieva, Valeriia Khylenko, Georg Groh

Despite the extensive amount of labeled datasets in the NLP text classification field, the persistent imbalance in data availability across various languages remains evident. Ukrainian, in particular, stands as a language that still can benefit from the continued refinement of cross-lingual methodologies. Due to our knowledge, there is a tremendous lack of Ukrainian corpora for typical text classification tasks. In this work, we leverage the state-of-the-art advances in NLP, exploring cross-lingual knowledge transfer methods avoiding manual data curation: large multilingual encoders and translation systems, LLMs, and language adapters. We test the approaches on three text classification tasks -- toxicity classification, formality classification, and natural language inference -- providing the "recipe" for the optimal setups.

📄 PDF Abstract BibTeX arXiv:2404.02043

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationNatural Language Inferencetext-classificationText ClassificationTransfer Learning

Similar Papers 제목 키워드 기반

SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses

2026-07-02 · Anna Chorna arxiv

Large Language Models are increasingly deployed in emotional-support contexts and crisis-related situations. Nevertheless, their cross-lingual abilities in these circumstances remain underexplored. Existing benchmarks em…

Toxicity Classification in Ukrainian

2024-04-27 · Daryna Dementieva, Valeriia Khylenko, Nikolay Babakov, Georg Groh

The task of toxicity detection is still a relevant task, especially in the context of safe and fair LMs development. Nevertheless, labeled binary toxicity classification corpora are not available for all languages, which…

ClassificationCross-Lingual TransferTransfer Learning

Some Notes on p(e)re-Reduplication in Bulgarian and Ukrainian: A Corpus-based Study

2022-09-01 · CLIB 2022 9 · Ivan Derzhanski, Olena Siruk

We present a comparative study of p(e)re-reduplication in Bulgarian and Ukrainian, based on material from a parallel corpus of bilingual texts. We analyse all occurrences found in the corpus of close sequences and conjun…

EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainian

2025-05-29 · Daryna Dementieva, Nikolay Babakov, Alexander Fraser

While Ukrainian NLP has seen progress in many texts processing tasks, emotion classification remains an underexplored area with no publicly available benchmark to date. In this work, we introduce EmoBench-UA, the first a…

Emotion Classification

Clustering Comparable Corpora of Russian and Ukrainian Academic Texts: Word Embeddings and Semantic Fingerprints

2016-04-18 · Andrey Kutuzov, Mikhail Kopotev, Tatyana Sviridenko, Lyubov Ivanova

We present our experience in applying distributional semantics (neural word embeddings) to the problem of representing and clustering documents in a bilingual comparable corpus. Our data is a collection of Russian and Uk…

ClusteringTranslationWord Embeddings