paper-with-me

홈 › Papers

COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval

2022-04-15 · CVPR 2022 1 · Haoyu Lu, Nanyi Fei, Yuqi Huo, Yizhao Gao, Zhiwu Lu, Ji-Rong Wen

Large-scale single-stream pre-training has shown dramatic performance in image-text retrieval. Regrettably, it faces low inference efficiency due to heavy attention layers. Recently, two-stream methods like CLIP and ALIGN with high inference efficiency have also shown promising performance, however, they only consider instance-level alignment between the two streams (thus there is still room for improvement). To overcome these limitations, we propose a novel COllaborative Two-Stream vision-language pretraining model termed COTS for image-text retrieval by enhancing cross-modal interaction. In addition to instance level alignment via momentum contrastive learning, we leverage two extra levels of cross-modal interactions in our COTS: (1) Token-level interaction - a masked visionlanguage modeling (MVLM) learning objective is devised without using a cross-stream network module, where variational autoencoder is imposed on the visual encoder to generate visual tokens for each image. (2) Task-level interaction - a KL-alignment learning objective is devised between text-to-image and image-to-text retrieval tasks, where the probability distribution per task is computed with the negative queues in momentum contrastive learning. Under a fair comparison setting, our COTS achieves the highest performance among all two-stream methods and comparable performance (but with 10,800X faster in inference) w.r.t. the latest single-stream methods. Importantly, our COTS is also applicable to text-to-video retrieval, yielding new state-ofthe-art on the widely-used MSR-VTT dataset.

📄 PDF Abstract BibTeX arXiv:2204.07441

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningCross-Modal RetrievalImage-text RetrievalImage to textImage-to-Text RetrievalRetrievalText RetrievalText to Video RetrievalVideo Retrieval

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Collaborative Tree Search for Enhancing Embodied Multi-Agent Collaboration

2025-01-01 · CVPR 2025 1 · Lizheng Zu, Lin Lin, Song Fu, Na Zhao 외

Embodied agents based on large language models (LLMs) face significant challenges in collaborative tasks, requiring effective communication and reasonable division of labor to ensure efficient and correct task comple…

Output Supervision Can Obfuscate the Chain of Thought

2025-10-11 · Jacob Drori, Luke Marks, Bryce Woodworth, Alex Cloud 외 arxiv

OpenAI (2025) showed that training against a chain of thought (CoT) monitor can cause obfuscated CoTs, which contain bad behavior the monitor cannot detect. They proposed to keep CoTs monitorable by training only against…

Whit’s the Richt Pairt o Speech: PoS tagging for Scots

2021-04-01 · EACL (VarDial) 2021 4 · Harm Lameris, Sara Stymne

In this paper we explore PoS tagging for the Scots language. Scots is spoken in Scotland and Northern Ireland, and is closely related to English. As no linguistically annotated Scots data were available, we manually PoS …

POSPOS TaggingTransfer Learning

Activation Steering for Chain-of-Thought Compression

2025-07-07 · Seyedarmin Azizi, Erfan Baghaei Potraghloo, Massoud Pedram

Large language models (LLMs) excel at complex reasoning when they include intermediate steps, known as "chains of thought" (CoTs). However, these rationales are often overly verbose, even for simple problems, leading to …

GSM8KMath

Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought

2025-05-18 · Hanlin Zhu, Shibo Hao, Zhiting Hu, Jiantao Jiao 외

Large Language Models (LLMs) have demonstrated remarkable performance in many applications, including challenging reasoning problems via chain-of-thoughts (CoTs) techniques that generate ``thinking tokens'' before answer…