paper-with-me

홈 › Papers

Rainbow Noise: Stress-Testing Multimodal Harmful-Meme Detectors on LGBTQ Content

2025-07-24 · Ran Tong, Songtao Wei, Jiaqi Liu, Lanruo Wang arxiv

Hateful memes aimed at LGBTQ\,+ communities often evade detection by tweaking either the caption, the image, or both. We build the first robustness benchmark for this setting, pairing four realistic caption attacks with three canonical image corruptions and testing all combinations on the PrideMM dataset. Two state-of-the-art detectors, MemeCLIP and MemeBLIP2, serve as case studies, and we introduce a lightweight \textbf{Text Denoising Adapter (TDA)} to enhance the latter's resilience. Across the grid, MemeCLIP degrades more gently, while MemeBLIP2 is particularly sensitive to the caption edits that disrupt its language processing. However, the addition of the TDA not only remedies this weakness but makes MemeBLIP2 the most robust model overall. Ablations reveal that all systems lean heavily on text, but architectural choices and pre-training data significantly impact robustness. Our benchmark exposes where current multimodal safety models crack and demonstrates that targeted, lightweight modules like the TDA offer a powerful path towards stronger defences.

📄 PDF Abstract BibTeX arXiv:2507.19551

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Novel Nongenomic Signaling by Glucocorticoid May Involve Changes to Liver Membrane Order in Rainbow Trout

2017-04-26

Stress-induced glucocorticoid elevation is a highly conserved response among vertebrates. This facilitates stress adaptation and the mode of action involves activation of the intracellular glucocorticoid receptor leading…

Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation

2026-04-18 · Huije Lee, Jisu Shin, Hoyun Song, Changgeon Ko 외 arxiv

Static benchmarks for harmful content detection face limitations in scalability and diversity, and may also be affected by contamination from web-scale pre-training corpora. To address these issues, we propose a framewor…

Stress-Testing Multimodal Foundation Models for Crystallographic Reasoning

2025-06-16 · Can Polat, Hasan Kurban, Erchin Serpedin, Mustafa Kurban

Evaluating foundation models for crystallographic reasoning requires benchmarks that isolate generalization behavior while enforcing physical constraints. This work introduces a multiscale multicrystal dataset with two p…

HallucinationSpatial Interpolation

Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability

2025-10-21 · Artur Zolkowski, Wen Xing, David Lindner, Florian Tramèr 외 arxiv

Recent findings suggest that misaligned models may exhibit deceptive behavior, raising concerns about output trustworthiness. Chain-of-thought (CoT) is a promising tool for alignment monitoring: when models articulate th…

Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique

2024-08-20 · Tej Deep Pala, Vernon Y. H. Toh, Rishabh Bhardwaj, Soujanya Poria

In today's era, where large language models (LLMs) are integrated into numerous real-world applications, ensuring their safety and robustness is crucial for responsible AI usage. Automated red-teaming methods play a key …

AI and SafetyDiversityRed TeamingSafety Alignment