paper-with-me

Papers

ShieldGemma: Generative AI Content Moderation Based on Gemma

2024-07-31 · Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, Joe Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush Kumar, Bhaktipriya Radharapu, Olivia Sturman, Oscar Wahltinez

We present ShieldGemma, a comprehensive suite of LLM-based safety content moderation models built upon Gemma2. These models provide robust, state-of-the-art predictions of safety risks across key harm types (sexually explicit, dangerous content, harassment, hate speech) in both user input and LLM-generated output. By evaluating on both public and internal benchmarks, we demonstrate superior performance compared to existing models, such as Llama Guard (+10.8\% AU-PRC on public benchmarks) and WildCard (+4.3\%). Additionally, we present a novel LLM-based data curation pipeline, adaptable to a variety of safety-related tasks and beyond. We have shown strong generalization performance for model trained mainly on synthetic data. By releasing ShieldGemma, we provide a valuable resource to the research community, advancing LLM safety and enabling the creation of more effective content moderation solutions for developers.

📄 PDF Abstract BibTeX arXiv:2407.21772

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

ShieldGemma 2: Robust and Tractable Image Content Moderation

2025-04-01 · Wenjun Zeng, Dana Kurniawan, Ryan Mullins, Yuchi Liu 외

We introduce ShieldGemma 2, a 4B parameter image content moderation model built on Gemma 3. This model provides robust safety risk predictions across the following key harm categories: Sexually Explicit, Violence \& Gore…

Image GenerationLanguage ModelingLanguage Modelling

KidsNanny: A Two-Stage Multimodal Content Moderation Pipeline Integrating Visual Classification, Object Detection, OCR, and Contextual Reasoning for Child Safety

2026-03-17 · Viraj Panchal, Tanmay Talsaniya, Parag Patel, Meet Patel arxiv

We present KidsNanny, a two-stage multimodal content moderation architecture for child safety. Stage 1 combines a vision transformer (ViT) with an object detector for visual screening (11.7 ms); outputs are routed as tex…

Object Detection

BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems

2026-06-12 · Leonhard Waibl, Felix Michalak, Hadrien Mariaccia arxiv

LLM supervision systems, namely input/output moderation filters and jailbreak detectors, are the primary safeguard against misuse in deployed AI applications, yet existing benchmarks are often vendor-biased, omit cost an…

AI Content Moderation in Therapy Conversations

2026-05-25 · Jiwon Kim, Claire Wang, Taeung Yoon, Sabelle Huang 외 arxiv

Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LLMs like ChaptGPT or Llama are often developed with content moderation …

A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness

2025-12-22 · Naseem Machlovi, Maryam Saleki, Ruhul Amin, Mohamed Rahouti 외 arxiv

As large language models (LLMs) become deeply embedded in daily life, the urgent need for safer moderation systems that distinguish between naive and harmful requests while upholding appropriate censorship boundaries has…

Adversarial Robustness