paper-with-me

홈 › Papers

Zero-shot Cross-lingual Content Filtering: Offensive Language and Hate Speech Detection

2021-04-01 · EACL (Hackashop) 2021 4 · Andraž Pelicon, Ravi Shekhar, Matej Martinc, Blaž Škrlj, Matthew Purver, Senja Pollak

We present a system for zero-shot cross-lingual offensive language and hate speech classification. The system was trained on English datasets and tested on a task of detecting hate speech and offensive social media content in a number of languages without any additional training. Experiments show an impressive ability of both models to generalize from English to other languages. There is however an expected gap in performance between the tested cross-lingual models and the monolingual models. The best performing model (offensive content classifier) is available online as a REST API.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Hate Speech Detection

Similar Papers 제목 키워드 기반

CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications

2025-08-03 · Raviraj Joshi, Rakesh Paul, Kanishk Singla, Anusha Kamath 외 arxiv

The increasing use of Large Language Models (LLMs) in agentic applications highlights the need for robust safety guard models. While content safety in English is well-studied, non-English languages lack similar advanceme…

Synthetic Data GenerationZero-shot GeneralizationCross-Lingual TransferMachine Translation

Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability

2023-05-17 · Eleftheria Briakou, Colin Cherry, George Foster

Large, multilingual language models exhibit surprisingly good zero- or few-shot machine translation capabilities, despite having never seen the intentionally-included translation examples provided to typical neural trans…

Language ModelingLanguage ModellingMachine TranslationTranslation

Zero and Few-Shot Localization of Task-Oriented Dialogue Agents with a Distilled Representation

2023-02-18 · Mehrad Moradshahi, Sina J. Semnani, Monica S. Lam

Task-oriented Dialogue (ToD) agents are mostly limited to a few widely-spoken languages, mainly due to the high cost of acquiring training data for each language. Existing low-cost approaches that rely on cross-lingual e…

Dialogue State TrackingMachine TranslationTranslation

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

2026-06-05 · Ahmer Tabassum, Sarfraz Ahmad, Hasan Iqbal, Owais Aijaz 외 arxiv

Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lacks a broad MMLU-style benchmark built from native educational sources. …

Medical Crossing: a Cross-lingual Evaluation of Clinical Entity Linking

2022-06-01 · LREC 2022 6 · Anton Alekseev, Zulfat Miftahutdinov, Elena Tutubalina, Artem Shelmanov 외

Medical data annotation requires highly qualified expertise. Despite the efforts devoted to medical entity linking in different languages, available data is very sparse in terms of both data volume and languages. In this…

Cross-Lingual TransferEntity LinkingTransfer LearningZero-Shot Cross-Lingual Transfer