paper-with-me

홈 › Papers

KIT-TIP-NLP at MultiPride: Continual Learning with Multilingual Foundation Model

2026-05-13 · Barathi Ganesh HB, Michal Ptaszynski, Rene Melendez, Juuso Eronen arxiv

This paper presents a multi-stage framework for detecting reclaimed slurs in multilingual social media discourse. It addresses the challenge of identifying reclamatory versus non-reclamatory usage of LGBTQ+-related slurs across English, Spanish, and Italian tweets. The framework handles three intertwined methodological challenges like data scarcity, class imbalance, and cross-linguistic variation in sentiment expression. It integrates data-driven model selection via cross-validation, semantic-preserving augmentation through back-translation, inductive transfer learning with dynamic epoch-level undersampling, and domain-specific knowledge injection via masked language modeling. Eight multilingual embedding models were evaluated systematically, with XLM-RoBERTa selected as the foundation model based on macro-averaged F1 score. Data augmentation via GPT-4o-mini back-translation to alternate languages effectively tripled the training corpus while preserving semantic content and class distribution ratios. The framework produces four final runs for the evaluation purposes where RUN 1 is inductive transfer learning with augmentation and undersampling, RUN 2 with masked language modeling pre-training, RUN 3 and RUN 4 are previous predictions refined via language-specific decision thresholds optimized via ROC analysis. Language-specific threshold refinement reveals that optimal decision boundaries vary significantly across languages. This reflects distributional differences in model confidence scores and linguistic variation in reclamatory language usage. The threshold-based optimization yields 2-5% absolute F1 improvement without requiring model retraining. The methodology is fully reproducible, with all code and experimental setup available at https://github.com/rbg-research/MultiPRIDE-Evalita-2026.

📄 PDF Abstract BibTeX arXiv:2605.13415

Code (0)

등록된 구현이 없습니다.

Tasks

Continual LearningTransfer LearningData Augmentation

Similar Papers 제목 키워드 기반

AIWizards at MULTIPRIDE: A Hierarchical Approach to Slur Reclamation Detection

2026-02-13 · Luca Tedeschini, Matteo Fasulo arxiv

Detecting reclaimed slurs represents a fundamental challenge for hate speech detection systems, as the same lexcal items can function either as abusive expressions or as in-group affirmations depending on social identity…

Hate Speech Detection

Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi

2026-03-03 · Shiza Fatimah, Aniket Sen, Sophia Falk, Florian Mai 외 arxiv

The dominance of large multilingual foundation models has widened linguistic inequalities in Natural Language Processing (NLP), often leaving low-resource languages underrepresented. This paper introduces LilMoo, a 0.6-b…

Continual Pretraining

CL-MASR: A Continual Learning Benchmark for Multilingual ASR

2023-10-25 · Luca Della Libera, Pooneh Mousavi, Salah Zaiem, Cem Subakan 외

Modern multilingual automatic speech recognition (ASR) systems like Whisper have made it possible to transcribe audio in multiple languages with a single model. However, current state-of-the-art ASR models are typically …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Continual Learningspeech-recognition+1

Challenger at MultiPRIDE: Is It Hate Speech or Reclaimed?

2026-05-31 · Hadi Bayrami Asl Tekanlou, Mahdi Bakhtiyarzadeh, Jafar Razmara arxiv

The spread of hate speech has become increasingly harmful in modern digital environments, particularly on social networking platforms. While recent advances have shown promising results in automatic hate speech detection…

Hate Speech Detection

Weight Factorization and Centralization for Continual Learning in Speech Recognition

2025-06-19 · Enes Yavuz Ugan, Ngoc-Quan Pham, Alexander Waibel

Modern neural network based speech recognition models are required to continually absorb new data without re-training the whole system, especially in downstream applications using foundation models, having no access to t…

Continual Learningspeech-recognitionSpeech Recognition