paper-with-me

홈 › Papers

A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval

2023-11-30 · Matthew Gwilliam, Michael Cogswell, Meng Ye, Karan Sikka, Abhinav Shrivastava, Ajay Divakaran

Existing long video retrieval systems are trained and tested in the paragraph-to-video retrieval regime, where every long video is described by a single long paragraph. This neglects the richness and variety of possible valid descriptions of a video, which could range anywhere from moment-by-moment detail to a single phrase summary. To provide a more thorough evaluation of the capabilities of long video retrieval systems, we propose a pipeline that leverages state-of-the-art large language models to carefully generate a diverse set of synthetic captions for long videos. We validate this pipeline's fidelity via rigorous human inspection. We use synthetic captions from this pipeline to perform a benchmark of a representative set of video language models using long video datasets, and show that the models struggle on shorter captions. We show that finetuning on this data can both mitigate these issues (+2.8% R@1 over SOTA on ActivityNet with diverse captions), and even improve performance on standard paragraph-to-video retrieval (+1.0% R@1 on ActivityNet). We also use synthetic data from our pipeline as query expansion in the zero-shot setting (+3.4% R@1 on ActivityNet). We derive insights by analyzing failure cases for retrieval with short captions. For data access and other details, please refer to our project website at https://mgwillia.github.io/10k-words.

📄 PDF Abstract BibTeX arXiv:2312.00115

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingRetrievalvalidVideo Retrieval

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

LUMA: A Benchmark Dataset for Learning from Uncertain and Multimodal Data

2024-06-14 · Grigor Bezirganyan, Sana Sellami, Laure Berti-ÉQuille, Sébastien Fournier

Multimodal Deep Learning enhances decision-making by integrating diverse information sources, such as texts, images, audio, and videos. To develop trustworthy multimodal approaches, it is essential to understand how unce…

BenchmarkingDecision MakingDiversityLanguage Modelling+4

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

2026-06-01 · Huiqiong Li, Jiayu Wang, Zhiting Mei, Anirudha Majumdar 외 arxiv

Video world models are increasingly used in robotic manipulation, yet existing benchmarks mostly evaluate them under valid, feasible, and safe instructions. We introduce RoboTrustBench, a benchmark for evaluating the tru…

Instruction Following

UniVTG: Towards Unified Video-Language Temporal Grounding

2023-07-31 · ICCV 2023 1 · Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick 외

Video Temporal Grounding (VTG), which aims to ground target clips from videos (such as consecutive intervals or disjoint shots) according to custom language queries (e.g., sentences or words), is key for video browsing o…

Highlight DetectionMoment RetrievalNatural Language Moment RetrievalRetrieval+1

A Closer Look at Debiased Temporal Sentence Grounding in Videos: Dataset, Metric, and Approach

2022-03-10 · Xiaohan Lan, Yitian Yuan, Xin Wang, Long Chen 외

Temporal Sentence Grounding in Videos (TSGV), which aims to ground a natural language sentence in an untrimmed video, has drawn widespread attention over the past few years. However, recent studies have found that curren…

BenchmarkingSentenceTemporal Sentence Grounding

A Video Is Not Worth a Thousand Words

2025-10-27 · Sam Pollard, Michael Wray arxiv

As we become increasingly dependent on vision language models (VLMs) to answer questions about the world around us, there is a significant amount of research devoted to increasing both the difficulty of video question an…

Video Question Answering