paper-with-me

Papers

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning

2026-04-05 · Tianle Chen, Deepti Ghadiyaram arxiv

As audio-visual multi-modal large language models (MLLMs) are increasingly deployed in safety-critical applications, understanding their vulnerabilities is crucial. To this end, we introduce Multi-Modal Typography, a systematic study examining how typographic attacks across multiple modalities adversely influence MLLMs. While prior work focuses narrowly on unimodal attacks, we expose the cross-modal fragility of MLLMs. We analyze the interactions between audio, visual, and text perturbations and reveal that coordinated multi-modal attack creates a significantly more potent threat than single-modality attacks (attack success rate = $83.43\%$ vs $34.93\%$).Our findings across multiple frontier MLLMs, tasks, and common-sense reasoning and content moderation benchmarks establishes multi-modal typography as a critical and underexplored attack strategy in multi-modal reasoning. Code and data will be publicly available.

📄 PDF Abstract BibTeX arXiv:2604.03995

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models

2025-04-07 · Justus Westerhoff, Erblina Purelku, Jakob Hackstein, Jonas Loos 외

Typographic attacks exploit the interplay between text and visual content in multimodal foundation models, causing misclassifications when misleading text is embedded within images. However, existing datasets are limited…

Benchmarking

Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Model

2024-02-29 · Hao Cheng, Erjia Xiao, Jindong Gu, Le Yang 외

Large Vision-Language Models (LVLMs) rely on vision encoders and Large Language Models (LLMs) to exhibit remarkable capabilities on various multi-modal tasks in the joint space of vision and language. However, the Typogr…

Language ModelingLanguage ModellingObject RecognitionZero-Shot Learning

Turning Adversaries into Allies: Reversing Typographic Attacks for Multimodal E-Commerce Product Retrieval

2025-11-07 · Janet Jenq, Hongda Shen arxiv

Multimodal product retrieval systems in e-commerce platforms rely on effectively combining visual and textual signals to improve search relevance and user experience. However, vision-language models such as CLIP are vuln…

Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP

2025-08-28 · Lorenz Hufe, Constantin Venhoff, Erblina Purelku, Maximilian Dreyer 외 arxiv

Typographic attacks exploit multi-modal systems by injecting text into images, leading to targeted misclassifications, malicious content generation and even Vision-Language Model jailbreaks. In this work, we analyze how …

AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents

2025-10-05 · Yanjie Li, Yiming Cao, Dong Wang, Bin Xiao arxiv

Multimodal agents built on large vision-language models (LVLMs) are increasingly deployed in open-world settings but remain highly vulnerable to prompt injection, especially through visual inputs. We introduce AgentTypo,…

Continual Learning