paper-with-me

홈 › Papers

The Cross-Domain Generalization Cost of Offensive Language Detection

2026-07-26 · Ruixing Ren, Junhui Zhao, Xiaoke Sun, Qiuping Li arxiv

Offensive language detection models generally suffer performance degradation when deployed across datasets and across languages, yet most existing studies stop at reporting this phenomenon and lack a systematic methodology for decomposing the causes of degradation into attributable components and quantifying the cost of remediation. This paper proposes a diagnosis and optimization framework composed of three coordinated technical components. First, a zero-shot transfer loss decomposition that separates the performance degradation from OLID to MLMA into two independently measurable components, namely dataset effect and language effect. Second, a controlled fine-tuning protocol that quantifies both adaptation efficiency and the hidden damage inflicted on the source task by comparing few shot learning curves under continued fine-tuning and cold-start starting points. Third, three joint training strategies incorpo rating temperature sampling and experience replay, which offer a controllable Pareto trade-off between improving multilingual capability and preserving source-task performance. Experiments built on this framework show that the dataset effect dominates the zero-shot transfer loss and substantially outweighs the language effect. Few-shot adaptation without a replay mechanism, though data-efficient, inflicts source task damage 4 to 9 times greater than that of the joint training strategies, and its damage magnitude is highly unstable. The three joint training strategies trade 3.2 to 4.1 percentage points of source-task performance for 8.1 to 42.6 percentage points of multilingual capability gain, forming a clear and controllable Pareto trade-off.

📄 PDF Abstract BibTeX arXiv:2607.23512

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Generalization

Similar Papers 제목 키워드 기반

A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection

2020-05-01 · LREC 2020 5 · Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Abdelali, Soon-gyo Jung 외

Access to social media often enables users to engage in conversation with limited accountability. This allows a user to share their opinions and ideology, especially regarding public content, occasionally adopting offens…

Cross-lingual Inductive Transfer to Detect Offensive Language

2020-07-07 · Kartikey Pant, Tanvi Dadu

With the growing use of social media and its availability, many instances of the use of offensive language have been observed across multiple languages and domains. This phenomenon has given rise to the growing need to d…

Language IdentificationPositionXLM-RZero-Shot Learning

Team Rouges at SemEval-2020 Task 12: Cross-lingual Inductive Transfer to Detect Offensive Language

2020-12-01 · SEMEVAL 2020 · Tanvi Dadu, Kartikey Pant

With the growing use of social media and its availability, many instances of the use of offensive language have been observed across multiple languages and domains. This phenomenon has given rise to the growing need to d…

Language IdentificationPositionXLM-RZero-Shot Learning

Cross-lingual Offensive Language Identification for Low Resource Languages: The Case of Marathi

2021-09-08 · RANLP 2021 9 · Saurabh Gaikwad, Tharindu Ranasinghe, Marcos Zampieri, Christopher M. Homan

The widespread presence of offensive language on social media motivated the development of systems capable of recognizing such content automatically. Apart from a few notable exceptions, most research on automatic offens…

Language IdentificationTransfer Learning

Developing Linguistic Patterns to Mitigate Inherent Human Bias in Offensive Language Detection

2023-12-04 · Toygar Tanyel, Besher Alkurdi, Serkan Ayvaz

With the proliferation of social media, there has been a sharp increase in offensive content, particularly targeting vulnerable groups, exacerbating social problems such as hatred, racism, and sexism. Detecting offensive…

Data AugmentationFairness