paper-with-me

홈 › Papers

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation

2026-04-09 · Wenli Zhang, Xianglong Shi, Sirui Zhao, Xinqi Chen, Guo Cheng, Yifan Xu, Tong Xu, Yong Liao arxiv

Diffusion-based audio-driven talking-head generation enables realistic portrait animation, but also introduces risks of misuse, such as fraud and misinformation. Existing protection methods are largely limited to a single modality, and neither image-only nor audio-only attacks can effectively suppress speech-driven facial dynamics. To address this gap, we propose SyncBreaker, a stage-aware multimodal protection framework that jointly perturbs portrait and audio inputs under modality-specific perceptual constraints. Our key contributions are twofold. First, for the image stream, we introduce nullifying supervision with Multi-Interval Sampling (MIS) across diffusion stages to steer the generation toward the static reference portrait by aggregating guidance from multiple denoising intervals. Second, for the audio stream, we propose Cross-Attention Fooling (CAF), which suppresses interval-specific audio-conditioned cross-attention responses. Both streams are optimized independently and combined at inference time to enable flexible deployment. We evaluate SyncBreaker in a white-box proactive protection setting. Extensive experiments demonstrate that SyncBreaker more effectively degrades lip synchronization and facial dynamics than strong single-modality baselines, while preserving input perceptual quality and remaining robust under purification. Code: https://github.com/kitty384/SyncBreaker.

📄 PDF Abstract BibTeX arXiv:2604.08405

Code (0)

등록된 구현이 없습니다.

Tasks

Talking Head Generation

Similar Papers 제목 키워드 기반

A Two-Stage Globally-Diverse Adversarial Attack for Vision-Language Pre-training Models

2026-01-18 · Wutao Chen, Huaqin Zou, Chen Wan, Lifeng Huang arxiv

Vision-language pre-training (VLP) models are vulnerable to adversarial examples, particularly in black-box scenarios. Existing multimodal attacks often suffer from limited perturbation diversity and unstable multi-stage…

Adversarial Attack

Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack

2025-05-28 · Juan Ren, Mark Dras, Usman Naseem

Large Vision-Language Models (LVLMs) have shown remarkable capabilities across a wide range of multimodal tasks. However, their integration of visual inputs introduces expanded attack surfaces, thereby exposing them to n…

Adversarial AttackSafety Alignment

Vulnerability-Aware Robust Multimodal Adversarial Training

2025-11-22 · Junrui Zhang, Xinyu Zhao, Jie Peng, Chenjie Wang 외 arxiv

Multimodal learning has shown significant superiority on various tasks by integrating multiple modalities. However, the interdependencies among modalities increase the susceptibility of multimodal models to adversarial a…

On the Robustness of Large Multimodal Models Against Image Adversarial Attacks

2023-12-06 · CVPR 2024 1 · Xuanming Cui, Alejandro Aparcedo, Young Kyun Jang, Ser-Nam Lim

Recent advances in instruction tuning have led to the development of State-of-the-Art Large Multimodal Models (LMMs). Given the novelty of these models, the impact of visual adversarial attacks on LMMs has not been thoro…

Image Captioningimage-classificationImage ClassificationVisual Question Answering (VQA)

When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models

2025-11-20 · Yuping Yan, Yuhan Xie, Yixin Zhang, Lingjuan Lyu 외 arxiv

Vision-Language-Action models (VLAs) have recently demonstrated remarkable progress in embodied environments, enabling robots to perceive, reason, and act through unified multimodal understanding. Despite their impressiv…

Semantic correspondenceAdversarial Robustness