paper-with-me

홈 › Papers

Representations of Text and Images Align From Layer One

2026-01-12 · Evžen Wybitul, Javier Rando, Florian Tramèr, Stanislav Fort arxiv

We show that for a variety of concepts in adapter-based vision-language models, the representations of their images and their text descriptions are meaningfully aligned from the very first layer. This contradicts the established view that such image-text alignment only appears in late layers. We show this using a new synthesis-based method inspired by DeepDream: given a textual concept such as "Jupiter", we extract its concept vector at a given layer, and then use optimisation to synthesise an image whose representation aligns with that vector. We apply our approach to hundreds of concepts across seven layers in Gemma 3, and find that the synthesised images often depict salient visual features of the targeted textual concepts: for example, already at layer 1, more than 50 % of images depict recognisable features of animals, activities, or seasons. Our method thus provides direct, constructive evidence of image-text alignment on a concept-by-concept and layer-by-layer basis. Unlike previous methods for measuring multimodal alignment, our approach is simple, fast, and does not require auxiliary models or datasets. It also offers a new path towards model interpretability, by providing a way to visualise a model's representation space by backtracing through its image processing components.

📄 PDF Abstract BibTeX arXiv:2601.08017

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Image Generation Based on Image Style Extraction

2025-10-01 · Shuochen Chang arxiv

Image generation based on text-to-image generation models is a task with practical application scenarios that fine-grained styles cannot be precisely described and controlled in natural language, while the guidance infor…

Text-to-Image Generation

Evaluating Image Editing with LLMs: A Comprehensive Benchmark and Intermediate-Layer Probing Approach

2026-03-20 · Shiqi Gao, Zitong Xu, Kang Fu, Huiyu Duan 외 arxiv

Evaluating text-guided image editing (TIE) methods remains a challenging problem, as reliable assessment should simultaneously consider perceptual quality, alignment with textual instructions, and preservation of origina…

Image Editing

Learning Shared Semantic Space with Correlation Alignment for Cross-modal Event Retrieval

2019-01-14 · Zhenguo Yang, Zehang Lin, Peipei Kang, Jianming Lv 외

In this paper, we propose to learn shared semantic space with correlation alignment (${S}^{3}CA$) for multimodal data representations, which aligns nonlinear correlations of multimodal data distributions in deep neural n…

ArticlesRetrieval

Neural Image Representations for Multi-Image Fusion and Layer Separation

2021-08-02 · Seonghyeon Nam, Marcus A. Brubaker, Michael S. Brown

We propose a framework for aligning and fusing multiple images into a single view using neural image representations (NIRs), also known as implicit or coordinate-based neural representations. Our framework targets burst …

Optical Flow Estimation

On convex decision regions in deep network representations

2023-05-26 · Lenka Tětková, Thea Brüsch, Teresa Karen Scheidt, Fabian Martin Mager 외

Current work on human-machine alignment aims at understanding machine-learned latent spaces and their correspondence to human representations. G{\"a}rdenfors' conceptual spaces is a prominent framework for understanding …

Few-Shot Learning