paper-with-me

홈 › Papers

V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models

2025-09-18 · Qidong Wang, Junjie Hu, Ming Jiang arxiv

Recent advances in causal interpretability have extended from language models to vision-language models (VLMs), seeking to reveal their internal mechanisms through input interventions. While textual interventions often target semantics, visual interventions typically rely on coarse pixel-level perturbations, limiting semantic insights on multimodal integration. In this study, we introduce V-SEAM, a novel framework that combines Visual Semantic Editing and Attention Modulating for causal interpretation of VLMs. V-SEAM enables concept-level visual manipulations and identifies attention heads with positive or negative contributions to predictions across three semantic levels: objects, attributes, and relationships. We observe that positive heads are often shared within the same semantic level but vary across levels, while negative heads tend to generalize broadly. Finally, we introduce an automatic method to modulate key head embeddings, demonstrating enhanced performance for both LLaVA and InstructBLIP across three diverse VQA benchmarks. Our data and code are released at: https://github.com/petergit1/V-SEAM.

📄 PDF Abstract BibTeX arXiv:2509.14837

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Conditioning Matters: Stabilizing Inversion and Attention in Diffusion Image Editing

2026-06-12 · Zheyuan Zhan, Hongchen Li, Can Wang, Yinfei Ma 외 arxiv

Inversion-based image editing offers flexible and training-free control but still struggles with inversion accuracy and the trade-off between editing fidelity and background preservation. While recent methods improve inv…

Image Editing

VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing

2025-02-24 · Xiangpeng Yang, Linchao Zhu, Hehe Fan, Yi Yang

Recent advancements in diffusion models have significantly improved video generation and editing capabilities. However, multi-grained video editing, which encompasses class-level, instance-level, and part-level modificat…

Video EditingVideo Generation

FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing

2023-10-09 · Yuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen 외

Text-to-video editing aims to edit the visual appearance of a source video conditional on textual prompts. A major challenge in this task is to ensure that all frames in the edited video are visually consistent. Most rec…

Optical Flow EstimationText-to-Video EditingVideo Editing

Bootstrapping Top-down Information for Self-modulating Slot Attention

2024-11-04 · Dongwon Kim, Seoyeon Kim, Suha Kwak

Object-centric learning (OCL) aims to learn representations of individual objects within visual scenes without manual supervision, facilitating efficient and effective visual reasoning. Traditional OCL methods primarily …

ObjectObject DiscoveryVisual Reasoning

SeamEdit: A Black-Box VLM-Agnostic Pipeline for Large-Image Semantic Editing

2026-06-11 · Xiangyu Lyu, Dan Lei arxiv

Semantic region editing for large images must satisfy two requirements at the same time: high generative quality and natural integration with surrounding content. Some related methods rely on white-box models and leave t…