paper-with-me

Papers

ShieldGemma 2: Robust and Tractable Image Content Moderation

2025-04-01 · Wenjun Zeng, Dana Kurniawan, Ryan Mullins, Yuchi Liu, Tamoghna Saha, Dirichi Ike-Njoku, Jindong Gu, Yiwen Song, Cai Xu, Jingjing Zhou, Aparna Joshi, Shravan Dheep, Mani Malek, Hamid Palangi, Joon Baek, Rick Pereira, Karthik Narasimhan

We introduce ShieldGemma 2, a 4B parameter image content moderation model built on Gemma 3. This model provides robust safety risk predictions across the following key harm categories: Sexually Explicit, Violence \& Gore, and Dangerous Content for synthetic images (e.g. output of any image generation model) and natural images (e.g. any image input to a Vision-Language Model). We evaluated on both internal and external benchmarks to demonstrate state-of-the-art performance compared to LlavaGuard \citep{helff2024llavaguard}, GPT-4o mini \citep{hurst2024gpt}, and the base Gemma 3 model \citep{gemma_2025} based on our policies. Additionally, we present a novel adversarial data generation pipeline which enables a controlled, diverse, and robust image generation. ShieldGemma 2 provides an open image moderation tool to advance multimodal safety and responsible AI development.

📄 PDF Abstract BibTeX arXiv:2504.01081

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

ShieldGemma: Generative AI Content Moderation Based on Gemma

2024-07-31 · Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran 외

We present ShieldGemma, a comprehensive suite of LLM-based safety content moderation models built upon Gemma2. These models provide robust, state-of-the-art predictions of safety risks across key harm types (sexually exp…

KidsNanny: A Two-Stage Multimodal Content Moderation Pipeline Integrating Visual Classification, Object Detection, OCR, and Contextual Reasoning for Child Safety

2026-03-17 · Viraj Panchal, Tanmay Talsaniya, Parag Patel, Meet Patel arxiv

We present KidsNanny, a two-stage multimodal content moderation architecture for child safety. Stage 1 combines a vision transformer (ViT) with an object detector for visual screening (11.7 ms); outputs are routed as tex…

Object Detection

BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems

2026-06-12 · Leonhard Waibl, Felix Michalak, Hadrien Mariaccia arxiv

LLM supervision systems, namely input/output moderation filters and jailbreak detectors, are the primary safeguard against misuse in deployed AI applications, yet existing benchmarks are often vendor-biased, omit cost an…

An Image is Worth a Thousand Toxic Words: A Metamorphic Testing Framework for Content Moderation Software

2023-08-18 · Wenxuan Wang, Jingyuan Huang, Jen-tse Huang, Chang Chen 외

The exponential growth of social media platforms has brought about a revolution in communication and content dissemination in human society. Nevertheless, these platforms are being increasingly misused to spread toxic co…

Validating Multimedia Content Moderation Software via Semantic Fusion

2023-05-23 · Wenxuan Wang, Jingyuan Huang, Chang Chen, Jiazhen Gu 외

The exponential growth of social media platforms, such as Facebook and TikTok, has revolutionized communication and content publication in human society. Users on these platforms can publish multimedia content that deliv…

Sentencesoftware testing