paper-with-me

Papers

LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content

2024-06-24 · Jessica Foo, Shaun Khoo

As large language models (LLMs) become increasingly prevalent in a wide variety of applications, concerns about the safety of their outputs have become more significant. Most efforts at safety-tuning or moderation today take on a predominantly Western-centric view of safety, especially for toxic, hateful, or violent speech. In this paper, we describe LionGuard, a Singapore-contextualized moderation classifier that can serve as guardrails against unsafe LLM outputs. When assessed on Singlish data, LionGuard outperforms existing widely-used moderation APIs, which are not finetuned for the Singapore context, by 14% (binary) and up to 51% (multi-label). Our work highlights the benefits of localization for moderation classifiers and presents a practical and scalable approach for low-resource languages.

📄 PDF Abstract BibTeX arXiv:2407.10995

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators

2025-07-21 · Leanne Tan, Gabriel Chua, Ziyu Ge, Roy Ka-Wei Lee arxiv

Modern moderation systems increasingly support multiple languages, but often fail to address localisation and low-resource variants - creating safety gaps in real-world deployments. Small models offer a potential alterna…

Contextualized Counterspeech: Strategies for Adaptation, Personalization, and Evaluation

2024-12-10 · Lorenzo Cima, Alessio Miaschi, Amaury Trujillo, Marco Avvenuti 외

AI-generated counterspeech offers a promising and scalable strategy to curb online toxicity through direct replies that promote civil discourse. However, current counterspeech is one-size-fits-all, lacking adaptation to …

Persuasiveness

DirectProbe: Studying Representations without Classifiers

2021-04-13 · NAACL 2021 4 · Yichu Zhou, Vivek Srikumar

Understanding how linguistic structures are encoded in contextualized embedding could help explain their impressive performance across NLP@. Existing approaches for probing them usually call for training classifiers and …

Efficient, Uncertainty-based Moderation of Neural Networks Text Classifiers

2022-04-04 · Findings (ACL) 2022 5 · Jakob Smedegaard Andersen, Walid Maalej

To maximize the accuracy and increase the overall acceptance of text classifiers, we propose a framework for the efficient, in-operation moderation of classifiers' output. Our framework focuses on use cases in which F1-s…

Benchmarking

A Holistic Approach to Undesired Content Detection in the Real World

2022-08-05 · Todor Markov, Chong Zhang, Sandhini Agarwal, Tyna Eloundou 외

We present a holistic approach to building a robust and useful natural language classification system for real-world content moderation. The success of such a system relies on a chain of carefully designed and executed s…

Active Learning