paper-with-me

Papers

Multimodal Function Vectors for Visual Relations

2025-10-02 · Shuhao Fu, Esther Goldberg, Ying Nian Wu, Hongjing Lu arxiv

Large Multimodal Models (LMMs) demonstrate impressive in-context learning abilities from few multimodal demonstrations, yet the internal mechanisms supporting such task learning remain opaque. Building on prior work of Large Language Models, we show that a small subset of attention heads in Large Multimodal Models is responsible for transmitting representations of visual relations. The activations of these attention heads, termed function vectors, can be extracted and manipulated to alter an LMM's performance on relational tasks. First, using synthetic and real image datasets, we apply causal mediation analysis to identify attention heads that strongly influence relational predictions, and extract multimodal function vectors that improve zero-shot accuracy at inference time. We further demonstrate that these multimodal function vectors can be fine-tuned with a modest amount of training data, while keeping LMM parameters frozen, to significantly outperform in-context learning baselines. Finally, we show that relation-specific function vectors can be linearly combined to solve analogy problems involving novel and untrained visual relations, highlighting the strong generalization ability of this approach. Through experiments on two LMMs, including OpenFlamingo and Qwen3-VL, our results show that these models encode visual relational knowledge within localized internal structures, which can be systematically extracted and optimized, thereby advancing our understanding of model modularity and enhancing control over relational reasoning in LMMs.

📄 PDF Abstract BibTeX arXiv:2510.02528

Code (0)

등록된 구현이 없습니다.

Tasks

Relational Reasoning

Similar Papers 제목 키워드 기반

Stochastic Neighbor Embedding of Multimodal Relational Data for Image-Text Simultaneous Visualization

2020-05-02 · Morihiro Mizutani, Akifumi Okuno, Geewook Kim, Hidetoshi Shimodaira

Multimodal relational data analysis has become of increasing importance in recent years, for exploring across different domains of data, such as images and their text tags obtained from social networking services (e.g., …

Graph Embedding

Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models

2025-05-20 · Woody Haosheng Gan, Deqing Fu, Julian Asilis, Ollie Liu 외

Steering methods have emerged as effective and targeted tools for guiding large language models' (LLMs) behavior without modifying their parameters. Multimodal large language models (MLLMs), however, do not currently enj…

Diversity

OPDR: Order-Preserving Dimension Reduction for Semantic Embedding of Multimodal Scientific Data

2024-08-15 · Chengyu Gong, Gefei Shen, Luanzheng Guo, Nathan Tallent 외

One of the most common operations in multimodal scientific data management is searching for the $k$ most similar items (or, $k$-nearest neighbors, KNN) from the database after being provided a new item. Although recent a…

Dimensionality Reduction

Multimodal Prediction based on Graph Representations

2019-12-21 · Icaro Cavalcante Dourado, Salvatore Tabbone, Ricardo da Silva Torres

This paper proposes a learning model, based on rank-fusion graphs, for general applicability in multimodal prediction tasks, such as multimodal regression and image classification. Rank-fusion graphs encode information f…

image-classificationImage ClassificationPredictionRetrieval

Deep Metric Learning using Similarities from Nonlinear Rank Approximations

2019-09-20 · Konstantin Schall, Kai Uwe Barthel, Nico Hezel, Klaus Jung

In recent years, deep metric learning has achieved promising results in learning high dimensional semantic feature embeddings where the spatial relationships of the feature vectors match the visual similarities of the im…

Metric LearningRetrieval