paper-with-me

홈 › Papers

Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models

2025-09-25 · Zoe Wanying He, Sean Trott, Meenakshi Khosla arxiv

Recent studies show that deep vision-only and language-only models--trained on disjoint modalities--nonetheless project their inputs into a partially aligned representational space. Yet we still lack a clear picture of where in each network this convergence emerges, what visual or linguistic cues support it, whether it captures human preferences in many-to-many image-text scenarios, and how aggregating exemplars of the same concept affects alignment. Here, we systematically investigate these questions. We find that alignment peaks in mid-to-late layers of both model types, reflecting a shift from modality-specific to conceptually shared representations. This alignment is robust to appearance-only changes but collapses when semantics are altered (e.g., object removal or word-order scrambling), highlighting that the shared code is truly semantic. Moving beyond the one-to-one image-caption paradigm, a forced-choice "Pick-a-Pic" task shows that human preferences for image-caption matches are mirrored in the embedding spaces across all vision-language model pairs. This pattern holds bidirectionally when multiple captions correspond to a single image, demonstrating that models capture fine-grained semantic distinctions akin to human judgments. Surprisingly, averaging embeddings across exemplars amplifies alignment rather than blurring detail. Together, our results demonstrate that unimodal networks converge on a shared semantic code that aligns with human judgments and strengthens with exemplar aggregation.

📄 PDF Abstract BibTeX arXiv:2509.20751

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Seeing without Pixels: Perception from Camera Trajectories

2025-11-26 · Zihui Xue, Kristen Grauman, Dima Damen, Andrew Zisserman 외 arxiv

Can one perceive a video's content without seeing its pixels, just from the camera trajectory-the path it carves through space? This paper is the first to systematically investigate this seemingly implausible question. T…

Camera Pose EstimationContrastive Learning

Seeing Through the Clouds: Cloud Gap Imputation with Prithvi Foundation Model

2024-04-30 · Denys Godwin, Hanxi Li, Michael Cecil, Hamed Alemohammad

Filling cloudy pixels in multispectral satellite imagery is essential for accurate data analysis and downstream applications, especially for tasks which require time series data. To address this issue, we compare the per…

Generative Adversarial NetworkImputationTime Series

Seeing in Words: Learning to Classify through Language Bottlenecks

2023-06-29 · Khalid Saifullah, Yuxin Wen, Jonas Geiping, Micah Goldblum 외

Neural networks for computer vision extract uninterpretable features despite achieving high accuracy on benchmarks. In contrast, humans can explain their predictions using succinct and intuitive descriptions. To incorpor…

Seeing Through Clouds in Satellite Images

2021-06-15 · Mingmin Zhao, Peder A. Olsen, Ranveer Chandra

This paper presents a neural-network-based solution to recover pixels occluded by clouds in satellite images. We leverage radio frequency (RF) signals in the ultra/super-high frequency band that penetrate clouds to help …

Cloud Removal

Domain-Specific Machine Translation to Translate Medicine Brochures in English to Sorani Kurdish

2025-01-23 · Mariam Shamal, Hossein Hassani

Access to Kurdish medicine brochures is limited, depriving Kurdish-speaking communities of critical health information. To address this problem, we developed a specialized Machine Translation (MT) model to translate Engl…

Machine TranslationSentenceTranslation