paper-with-me

홈 › Papers

VeRVE: Versatile Retrieval for Videos via Unified Embeddings

2026-01-17 · Shaunak Halbe, Bhagyashree Puranik, Jayakrishnan Unnikrishnan, Kushan Thakkar, Vimal Bhat, Toufiq Parag arxiv

Modern video retrieval systems are expected to handle diverse tasks ranging from corpus-level retrieval, fine-grained moment localization to flexible multimodal querying. Specialized architectures achieve strong retrieval performance by training modality-specific encoders on massive datasets, but they lack the ability to process composed multimodal queries. In contrast, multimodal LLM (MLLM)-based methods support rich multimodal search but their retrieval performance remains well below that of specialized systems. We present VeRVE, an MLLM-based versatile video retrieval framework that integrates corpus and moment-level retrieval capabilities while accommodating composed multimodal queries within a single architecture. We use contrastive alignment of visual and textual embeddings generated using a shared MLLM backbone to facilitate efficient embedding-based candidate search. Our embedding model, trained efficiently using low-rank adaptation (LoRA) on 700K paired visual-text data samples, surpasses other MLLM-based methods on zero-shot video retrieval tasks. Additionally, we demonstrate that the same model can be adapted without further training to achieve competitive results on zero-shot moment retrieval, and state of the art results for zero-shot composed video retrieval. With additional training for reranking candidates identified in the embedding-based search, our model substantially outperforms existing MLLM-based retrieval systems and achieves retrieval performance comparable to state of the art specialized models.

📄 PDF Abstract BibTeX arXiv:2601.12193

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot Moment RetrievalZero-Shot Video Retrieval

Similar Papers 제목 키워드 기반

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

2025-09-26 · Changli Tang, Qinfan Xiao, Ke Mei, Tianyi Wang 외 arxiv

While embeddings from multimodal large language models (LLMs) excel as general-purpose representations, their application to dynamic modalities like audio and video remains underexplored. We introduce WAVE (\textbf{u}nif…

Cross-Modal RetrievalQuestion Answering

OmniSearchSage: Multi-Task Multi-Entity Embeddings for Pinterest Search

2024-04-25 · Prabhat Agarwal, Minhazul Islam Sk, Nikil Pancha, Kurchi Subhra Hazra 외

In this paper, we present OmniSearchSage, a versatile and scalable system for understanding search queries, pins, and products for Pinterest search. We jointly learn a unified query embedding coupled with pin and product…

Entity EmbeddingsImage CaptioningMulti-Task Learning

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

2025-07-07 · Rui Meng, Ziyan Jiang, Ye Liu, Mingyi Su 외 arxiv

Multimodal embedding models have been crucial in enabling various downstream tasks such as semantic similarity, information retrieval, and clustering over different modalities. However, existing multimodal embeddings lik…

Video Question AnsweringRepresentation LearningInformation RetrievalVideo Classification

VERVE: Template-based ReflectiVE Rewriting for MotiVational IntErviewing

2023-11-14 · Do June Min, Verónica Pérez-Rosas, Kenneth Resnicow, Rada Mihalcea

Reflective listening is a fundamental skill that counselors must acquire to achieve proficiency in motivational interviewing (MI). It involves responding in a manner that acknowledges and explores the meaning of what the…

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset

2026-05-19 · Haojun Chen, Haoyang He, Chengming Xu, Qingdong He 외 arxiv

Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the rapid development of imaging technology, the demand for Ultra-High-…

Image Generation