paper-with-me

홈 › Papers

Do Neural Network Cross-Modal Mappings Really Bridge Modalities?

2018-05-19 · ACL 2018 7 · Guillem Collell, Marie-Francine Moens

Feed-forward networks are widely used in cross-modal applications to bridge modalities by mapping distributed vectors of one modality to the other, or to a shared space. The predicted vectors are then used to perform e.g., retrieval or labeling. Thus, the success of the whole system relies on the ability of the mapping to make the neighborhood structure (i.e., the pairwise similarities) of the predicted vectors akin to that of the target vectors. However, whether this is achieved has not been investigated yet. Here, we propose a new similarity measure and two ad hoc experiments to shed light on this issue. In three cross-modal benchmarks we learn a large number of language-to-vision and vision-to-language neural network mappings (up to five layers) using a rich diversity of image and text features and loss functions. Our results reveal that, surprisingly, the neighborhood structure of the predicted vectors consistently resembles more that of the input vectors than that of the target vectors. In a second experiment, we further show that untrained nets do not significantly disrupt the neighborhood (i.e., semantic) structure of the input vectors.

📄 PDF Abstract BibTeX arXiv:1805.07616

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityRetrieval

Similar Papers 제목 키워드 기반

EarthBridge: A Solution for 4th Multi-modal Aerial View Image Challenge Translation Track

2026-03-06 · Zhenyuan Chen, Guanyuan Shen, Feng Zhang arxiv

Cross-modal image-to-image translation among Electro-Optical (EO), Infrared (IR), and Synthetic Aperture Radar (SAR) sensors is essential for comprehensive multi-modal aerial-view analysis. However, translating between t…

Image-to-Image TranslationContrastive Learning

DM2C: Deep Mixed-Modal Clustering

2019-12-01 · NeurIPS 2019 12 · Yangbangyan Jiang, Qianqian Xu, Zhiyong Yang, Xiaochun Cao 외

Data exhibited with multiple modalities are ubiquitous in real-world clustering tasks. Most existing methods, however, pose a strong assumption that the pairing information for modalities is available for all instances. …

AllClustering

Diffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image Captioning

2025-01-01 · CVPR 2025 1 · Jeong Ryong Lee, Yejee Shin, Geonhui Son, Dosik Hwang

The modality gap between vision and text embeddings in CLIP presents a significant challenge for zero-shot image captioning, limiting effective cross-modal representation. Traditional approaches, such as noise inject…

cross-modal alignmentDenoisingImage Captioning

Activation Space Interventions Can Be Transferred Between Large Language Models

2025-03-06 · Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Michael Lan 외

The study of representation universality in AI models reveals growing convergence across domains, modalities, and architectures. However, the practical applications of representation universality remain largely unexplore…

BioX-Bridge: Model Bridging for Unsupervised Cross-Modal Knowledge Transfer across Biosignals

2025-10-02 · Chenqi Li, Yu Liu, Timothy Denison, Tingting Zhu arxiv

Biosignals offer valuable insights into the physiological states of the human body. Although biosignal modalities differ in functionality, signal fidelity, sensor comfort, and cost, they are often intercorrelated, reflec…

Knowledge Distillation