paper-with-me

Papers

MATE: Meet At The Embedding -- Connecting Images with Long Texts

2024-06-26 · Young Kyun Jang, Junmo Kang, Yong Jae Lee, Donghyun Kim

While advancements in Vision Language Models (VLMs) have significantly improved the alignment of visual and textual data, these models primarily focus on aligning images with short descriptive captions. This focus limits their ability to handle complex text interactions, particularly with longer texts such as lengthy captions or documents, which have not been extensively explored yet. In this paper, we introduce Meet At The Embedding (MATE), a novel approach that combines the capabilities of VLMs with Large Language Models (LLMs) to overcome this challenge without the need for additional image-long text pairs. Specifically, we replace the text encoder of the VLM with a pretrained LLM-based encoder that excels in understanding long texts. To bridge the gap between VLM and LLM, MATE incorporates a projection module that is trained in a multi-stage manner. It starts by aligning the embeddings from the VLM text encoder with those from the LLM using extensive text pairs. This module is then employed to seamlessly align image embeddings closely with LLM embeddings. We propose two new cross-modal retrieval benchmarks to assess the task of connecting images with long texts (lengthy captions / documents). Extensive experimental results demonstrate that MATE effectively connects images with long texts, uncovering diverse semantic relationships.

📄 PDF Abstract BibTeX arXiv:2407.09541

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalDescriptive

Methods 이 논문이 사용한 방법론

Focus 설명 없음
MATE MATE is a Transformer architecture designed to model the structure of web tables. It uses sparse attention in a way that…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Connecting Look and Feel: Associating the visual and tactile properties of physical materials

2017-04-12 · CVPR 2017 7 · Wenzhen Yuan, Shaoxiong Wang, Siyuan Dong, Edward Adelson

For machines to interact with the physical world, they must understand the physical properties of objects and materials they encounter. We use fabrics as an example of a deformable material with a rich set of mechanical …

Theory and Approximate Solvers for Branched Optimal Transport with Multiple Sources

2022-10-14 · Peter Lippmann, Enrique Fita Sanmartín, Fred A. Hamprecht

Branched Optimal Transport (BOT) is a generalization of optimal transport in which transportation costs along an edge are subadditive. This subadditivity models an increase in transport efficiency when shipping mass alon…

Combinatorial Optimization

Understanding mobility in networks: A node embedding approach

2021-11-11 · Matheus F. C. Barros, Carlos H. G. Ferreira, Bruno Pereira dos Santos, Lourenço A. P. Júnior 외

Motivated by the growing number of mobile devices capable of connecting and exchanging messages, we propose a methodology aiming to model and analyze node mobility in networks. We note that many existing solutions in the…

Specificity

Connecting NeRFs, Images, and Text

2024-04-11 · Francesco Ballerini, Pierluigi Zama Ramirez, Roberto Mirabella, Samuele Salti 외

Neural Radiance Fields (NeRFs) have emerged as a standard framework for representing 3D scenes and objects, introducing a novel data type for information exchange and storage. Concurrently, significant progress has been …

NeRFRepresentation LearningRetrievalzero-shot-classification+1

Frame-wise and overlap-robust speaker embeddings for meeting diarization

2023-06-01 · Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă, Rama Doddipatla 외

Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible spea…