paper-with-me

홈 › Papers

Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum

2025-10-31 · Zhuoning Guo, Mingxin Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Xiaowen Chu arxiv

The prevailing video retrieval paradigm is structurally misaligned, as narrow benchmarks incentivize correspondingly limited data and single-task training. Therefore, universal capability is suppressed due to the absence of a diagnostic evaluation that defines and demands multi-dimensional generalization. To break this cycle, we introduce a framework built on the co-design of evaluation, data, and modeling. First, we establish the Universal Video Retrieval Benchmark (UVRB), a suite of 16 datasets designed not only to measure performance but also to diagnose critical capability gaps across tasks and domains. Second, guided by UVRB's diagnostics, we introduce a scalable synthesis workflow that generates 1.55 million high-quality pairs to populate the semantic space required for universality. Finally, we devise the Modality Pyramid, a curriculum that trains our General Video Embedder (GVE) by explicitly leveraging the latent interconnections within our diverse data. Extensive experiments show GVE achieves state-of-the-art zero-shot generalization on UVRB. In particular, our analysis reveals that popular benchmarks are poor predictors of general ability and that partially relevant retrieval is a dominant but overlooked scenario. Overall, our co-designed framework provides a practical path to escape the limited scope and advance toward truly universal video retrieval.

📄 PDF Abstract BibTeX arXiv:2510.27571

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationVideo Retrieval

Similar Papers 제목 키워드 기반

VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval

2026-02-08 · Issar Tzachor, Dvir Samuel, Rami Ben-Ari arxiv

Recent studies have adapted generative Multimodal Large Language Models (MLLMs) into embedding extractors for vision tasks, typically through fine-tuning to produce universal representations. However, their performance o…

Video-Text RetrievalVideo Retrieval

MuMUR : Multilingual Multimodal Universal Retrieval

2022-08-24 · Avinash Madasu, Estelle Aflalo, Gabriela Ben Melech Stan, Shachar Rosenman 외

Multi-modal retrieval has seen tremendous progress with the development of vision-language models. However, further improving these models require additional labelled data which is a huge manual effort. In this paper, we…

Image RetrievalMachine TranslationRetrievalTransfer Learning+1

Universal Adversarial Head: Practical Protection against Video Data Leakage

2021-06-18 · ICML Workshop AML 2021 7 · Jiawang Bai, Bin Chen, Dongxian Wu, Chaoning Zhang 외

While online video sharing becomes more popular, it also causes unconscious leakage of personal information in the video retrieval systems like deep hashing. An adversary can collect users' private information from the v…

Deep HashingRetrievalVideo Retrieval

ViLL-E: Video LLM Embeddings for Retrieval

2026-04-13 · Rohit Gupta, Jayakrishnan Unnikrishnan, Fan Fei, Sheng Liu 외 arxiv

Video Large Language Models (VideoLLMs) excel at video understanding tasks where outputs are textual, such as Video Question Answering and Video Captioning. However, they underperform specialized embedding-based models i…

Video Question AnsweringContrastive LearningMoment RetrievalVideo Captioning

U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs

2025-07-20 · Xiaojie Li, Chu Li, Shi-Zhe Chen, Xi Chen arxiv

Universal multimodal retrieval (UMR), which aims to address complex retrieval tasks where both queries and candidates span diverse modalities, has been significantly advanced by the emergence of MLLMs. While state-of-the…

Contrastive LearningImage RetrievalVideo Retrieval