paper-with-me

홈 › Papers

Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos

2025-07-03 · Zecheng Zhao, Selena Song, Tong Chen, Zhi Chen, Shazia Sadiq, Yadan Luo arxiv

Text-to-video (T2V) synthesis has advanced rapidly, yet current evaluation metrics primarily capture visual quality and temporal consistency, offering limited insight into how synthetic videos perform in downstream tasks such as text-to-video retrieval (TVR). In this work, we introduce SynTVA, a new dataset and benchmark designed to evaluate the utility of synthetic videos for building retrieval models. Based on 800 diverse user queries derived from MSRVTT training split, we generate synthetic videos using state-of-the-art T2V models and annotate each video-text pair along four key semantic alignment dimensions: Object \& Scene, Action, Attribute, and Prompt Fidelity. Our evaluation framework correlates general video quality assessment (VQA) metrics with these alignment scores, and examines their predictive power for downstream TVR performance. To explore pathways of scaling up, we further develop an Auto-Evaluator to estimate alignment quality from existing metrics. Beyond benchmarking, our results show that SynTVA is a valuable asset for dataset augmentation, enabling the selection of high-utility synthetic samples that measurably improve TVR outcomes. Project page and dataset can be found at https://jasoncodemaker.github.io/SynTVA/.

📄 PDF Abstract BibTeX arXiv:2507.02316

Code (0)

등록된 구현이 없습니다.

Tasks

Video Quality AssessmentVideo Retrieval

Similar Papers 제목 키워드 기반

From Third Person to First Person: Dataset and Baselines for Synthesis and Retrieval

2018-12-01 · Mohamed Elfeki, Krishna Regmi, Shervin Ardeshir, Ali Borji

First-person (egocentric) and third person (exocentric) videos are drastically different in nature. The relationship between these two views have been studied in recent years, however, it has yet to be fully explored. In…

Domain AdaptationGenerative Adversarial NetworkOptical Flow EstimationRetrieval

Retrieval-Augmented Egocentric Video Captioning

2024-01-01 · CVPR 2024 1 · Jilan Xu, Yifei HUANG, Junlin Hou, Guo Chen 외

Understanding human actions from videos of first-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while overlooking the potential benefit of explo…

Representation LearningRetrievalVideo Captioning

MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval

2024-10-15 · CVPR 2025 1 · Reno Kriz, Kate Sanders, David Etter, Kenton Murray 외

Efficiently retrieving and synthesizing information from large-scale multimodal collections has become a critical challenge. However, existing video retrieval datasets suffer from scope limitations, primarily focusing on…

DescriptiveRetrievalVideo Retrieval

EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark

2025-10-07 · Deheng Zhang, Yuqian Fu, Runyi Yang, Yang Miao 외 arxiv

Most existing benchmarks for understanding egocentric vision focus primarily on daytime scenarios, overlooking the low-light conditions that are inevitable in real-world applications. To investigate this gap, we present …

Visual Question AnsweringDepth Estimation

From My View to Yours: Ego-Augmented Learning in Large Vision Language Models for Understanding Exocentric Daily Living Activities

2025-01-10 · Dominick Reilly, Manish Kumar Govind, Le Xue, Srijan Das

Large Vision Language Models (LVLMs) have demonstrated impressive capabilities in video understanding, yet their adoption for Activities of Daily Living (ADL) remains limited by their inability to capture fine-grained in…

Human-Object Interaction DetectionKnowledge DistillationVideo Understanding