paper-with-me

Papers

AngleRoCL: Angle-Robust Concept Learning for Physically View-Invariant T2I Adversarial Patches

2025-06-11 · Wenjun Ji, Yuxiang Fu, Luyang Ying, Deng-Ping Fan, Yuyi Wang, Ming-Ming Cheng, Ivor Tsang, Qing Guo

Cutting-edge works have demonstrated that text-to-image (T2I) diffusion models can generate adversarial patches that mislead state-of-the-art object detectors in the physical world, revealing detectors' vulnerabilities and risks. However, these methods neglect the T2I patches' attack effectiveness when observed from different views in the physical world (i.e., angle robustness of the T2I adversarial patches). In this paper, we study the angle robustness of T2I adversarial patches comprehensively, revealing their angle-robust issues, demonstrating that texts affect the angle robustness of generated patches significantly, and task-specific linguistic instructions fail to enhance the angle robustness. Motivated by the studies, we introduce Angle-Robust Concept Learning (AngleRoCL), a simple and flexible approach that learns a generalizable concept (i.e., text embeddings in implementation) representing the capability of generating angle-robust patches. The learned concept can be incorporated into textual prompts and guides T2I models to generate patches with their attack effectiveness inherently resistant to viewpoint variations. Through extensive simulation and physical-world experiments on five SOTA detectors across multiple views, we demonstrate that AngleRoCL significantly enhances the angle robustness of T2I adversarial patches compared to baseline methods. Our patches maintain high attack success rates even under challenging viewing conditions, with over 50% average relative improvement in attack effectiveness across multiple angles. This research advances the understanding of physically angle-robust patches and provides insights into the relationship between textual concepts and physical properties in T2I-generated contents.

📄 PDF Abstract BibTeX arXiv:2506.09538

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

HeadLighter: Disentangling Illumination in Generative 3D Gaussian Heads via Lightstage Captures

2026-01-05 · Yating Wang, Yuan Sun, Xuan Wang, Ran Yi 외 arxiv

Recent 3D-aware head generative models based on 3D Gaussian Splatting achieve real-time, photorealistic and view-consistent head synthesis. However, a fundamental limitation persists: the deep entanglement of illuminatio…

PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning

2026-08-31 · Youngchae Chee, Hosu Lee, Sungjune Park, Junho Kim 외 arxiv

Cross-view video representation learning aims to capture viewpoint-invariant action semantics despite substantial appearance changes across egocentric and exocentric videos. However, existing methods encode each video as…

Representation Learning

Physically Disentangled Representations

2022-04-11 · Tzofi Klinghoffer, Kushagra Tiwary, Arkadiusz Balata, Vivek Sharma 외

State-of-the-art methods in generative representation learning yield semantic disentanglement, but typically do not consider physical scene parameters, such as geometry, albedo, lighting, or camera. We posit that inverse…

AttributeClassificationDisentanglementEmotion Recognition+2

Dense-View GEIs Set: View Space Covering for Gait Recognition based on Dense-View GAN

2020-09-26 · Rijun Liao, Weizhi An, Shiqi Yu, Zhu Li 외

Gait recognition has proven to be effective for long-distance human recognition. But view variance of gait features would change human appearance greatly and reduce its performance. Most existing gait datasets usually co…

Gait Recognition

View Invariant Learning for Vision-Language Navigation in Continuous Environments

2025-07-05 · Josh Qixuan Sun, Huaiyuan Weng, Xiaoying Xing, Chul Min Yeum 외 arxiv

Vision-Language Navigation in Continuous Environments (VLNCE), where an agent follows instructions and moves freely to reach a destination, is a key research problem in embodied AI. However, most existing approaches are …

Vision-Language NavigationContrastive Learning