paper-with-me

홈 › Papers

Security Tensors as a Cross-Modal Bridge: Extending Text-Aligned Safety to Vision in LVLM

2025-07-28 · Shen Li, Liuyi Yao, Wujia Niu, Lan Zhang, Yaliang Li arxiv

Large visual-language models (LVLMs) integrate aligned large language models (LLMs) with visual modules to process multimodal inputs. However, the safety mechanisms developed for text-based LLMs do not naturally extend to visual modalities, leaving LVLMs vulnerable to harmful image inputs. To address this cross-modal safety gap, we introduce security tensors - trainable input vectors applied during inference through either the textual or visual modality. These tensors transfer textual safety alignment to visual processing without modifying the model's parameters. They are optimized using a curated dataset containing (i) malicious image-text pairs requiring rejection, (ii) contrastive benign pairs with text structurally similar to malicious queries, with the purpose of being contrastive examples to guide visual reliance, and (iii) general benign samples preserving model functionality. Experimental results demonstrate that both textual and visual security tensors significantly enhance LVLMs' ability to reject diverse harmful visual inputs while maintaining near-identical performance on benign tasks. Further internal analysis towards hidden-layer representations reveals that security tensors successfully activate the language module's textual "safety layers" in visual inputs, thereby effectively extending text-based safety to the visual modality.

📄 PDF Abstract BibTeX arXiv:2507.20994

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-way Graph Signal Processing on Tensors: Integrative analysis of irregular geometries

2020-06-30 · Jay S. Stanley III, Eric C. Chi, Gal Mishne

Graph signal processing (GSP) is an important methodology for studying data residing on irregular structures. As acquired data is increasingly taking the form of multi-way tensors, new signal processing tools are needed …

Convex recovery of tensors using nuclear norm penalization

2015-06-08 · Stephane Chretien, Tianwen Wei

The subdifferential of convex functions of the singular spectrum of real matrices has been widely studied in matrix analysis, optimization and automatic control theory. Convex analysis and optimization over spaces of ten…

Deep Perceptual Mapping for Cross-Modal Face Recognition

2016-01-20 · M. Saquib Sarfraz, Rainer Stiefelhagen

Cross modal face matching between the thermal and visible spectrum is a much desired capability for night-time surveillance and security applications. Due to a very large modality gap, thermal-to-visible face recognition…

Face Recognition

Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge

2025-10-23 · Nimrod Berman, Omkar Joglekar, Eitan Kosman, Dotan Di Castro 외 arxiv

Recent advances in generative modeling have positioned diffusion models as state-of-the-art tools for sampling from complex data distributions. While these models have shown remarkable success across single-modality doma…

Image Super-Resolution

On the Geometry of Learned Representations in Event-Based Multi-Modal Egomotion Estimation

2026-07-17 · Stefano Silvestrini, Michele Ceresoli arxiv

Classical approaches to event-based egomotion estimation, including those adopted by the top-performing teams of the ELOPE challenge, rely on geometric optimization frameworks such as contrast maximization, homography es…

Homography Estimation