paper-with-me

Papers

UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages

2025-03-29 · Himanshu Beniwal, Reddybathuni Venkat, Rohit Kumar, Birudugadda Srivibhav, Daksh Jain, Pavan Doddi, Eshwar Dhande, Adithya Ananth, Kuldeep, Heer Kubadia, Pratham Sharda, Mayank Singh

This work introduces UnityAI-Guard, a framework for binary toxicity classification targeting low-resource Indian languages. While existing systems predominantly cater to high-resource languages, UnityAI-Guard addresses this critical gap by developing state-of-the-art models for identifying toxic content across diverse Brahmic/Indic scripts. Our approach achieves an impressive average F1-score of 84.23% across seven languages, leveraging a dataset of 888k training instances and 35k manually verified test instances. By advancing multilingual content moderation for linguistically diverse regions, UnityAI-Guard also provides public API access to foster broader adoption and application.

📄 PDF Abstract BibTeX arXiv:2503.23088

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive Perturbations

2026-03-16 · Hankun Kang, Xin Miao, Jianhao Chen, Jintao Wen 외 arxiv

Toxicity detection mitigates the dissemination of toxic content (e.g., hateful comments, posts, and messages within online social actions) to safeguard a healthy online social environment. However, malicious users persis…

Continual Learning

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

2026-05-28 · Ihor Stepanov, Aleksandr Smechov arxiv

Real-time safety filtering for large language model (LLM) applications requires classifiers that can detect unsafe prompts, toxic language, jailbreak attempts, and unsafe responses without the cost profile of large guard…

A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models

2026-06-24 · Soham Dan, Himanshu Beniwal, Thomas Hartvigsen arxiv

Large language models (LLMs) are increasingly deployed across languages, but their safety behavior remains uneven across linguistic and cultural contexts. This survey synthesizes work on toxicity detection and detoxifica…

DriftGuard: Safety-Aware Multi-Monitor Detection and Selective Adaptation for Evolving Toxicity Moderation

2026-06-27 · Yuting Xin, Hanyu Cai, Binqi Shen, Lier Jin 외 arxiv

Automated toxicity moderation systems operate in dynamic online environments where harmful behavior evolves through coded language, shifting targets, and strategic adaptation to enforcement. Existing drift detection meth…

Do You Feel Comfortable? Detecting Hidden Conversational Escalation in AI Chatbots

2025-12-05 · Jihyung Park, Saleh Afroogh, David Atkinson, Junfeng Jiao arxiv

Large Language Models (LLM) are increasingly integrated into everyday interactions, serving not only as information assistants but also as emotional companions. Even in the absence of explicit toxicity, repeated emotiona…