paper-with-me

홈 › Papers

ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search

2026-01-30 · Tao Yu, Haopeng Jin, Hao Wang, Shenghua Chai, Yujia Yang, Junhao Gong, Jiaming Guo, Minghui Zhang, Xinlong Chen, Zhenghao Zhang, Yuxuan Zhou, Yufei Xiong, Shanbin Zhang, Jiabing Yang, Hongzhu Yi, Xinming Wang, Cheng Zhong, Xiao Ma, Zhang Zhang, Yan Huang, Liang Wang arxiv

In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multimodal settings. Open-domain video shot retrieval, which involves richer temporal structure and more complex semantics, still lacks systematic benchmarks and analysis. To fill this gap, we introduce ShotFinder, a benchmark that formalizes editing requirements as keyframe-oriented shot descriptions and introduces five types of controllable single-factor constraints: Temporal order, Color, Visual style, Audio, and Resolution. We curate 1,210 high-quality samples from YouTube across 20 thematic categories, using large models for generation with human verification. Based on the benchmark, we propose ShotFinder, a text-driven three-stage retrieval and localization pipeline: (1) query expansion via video imagination, (2) candidate video retrieval with a search engine, and (3) description-guided temporal localization. Experiments on multiple closed-source and open-source models reveal a significant gap to human performance, with clear imbalance across constraints: temporal localization is relatively tractable, while color and visual style remain major challenges. These results reveal that open-domain video shot retrieval is still a critical capability that multimodal large models have yet to overcome.

📄 PDF Abstract BibTeX arXiv:2601.23232

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalVideo Retrieval

Similar Papers 제목 키워드 기반

A Survey of Emerging Approaches and Advances in Video Generation

2024-11-09 · HAL Preprint 2024 11 · Elnaz Soleimani, Ghazaleh Khodabandelou

The field of AI-driven video generation is evolving rapidly, with remarkable advancements achieved over the past two years. These developments have markedly enhanced the ability to transform human imagination into realis…

Image to Video GenerationLanguage ModelingLanguage ModellingLarge Language Model+5

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning

2026-05-31 · Xuchen Liu, Jiawei Huang, Shihao Xia, Bingxi Liu 외 arxiv

Vision-language navigation (VLN) for UAVs demands grounding free-form instructions into 6-DoF flight under partial observability. While Vision-Language-Action (VLA) models excel at semantic reasoning, they suffer from br…

Vision-Language Navigation

StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement

2026-05-29 · Junwon Seo, Sushant Veer, Ran Tian, Wenhao Ding 외 arxiv

Video world models (WMs) have shown promise for policy evaluation and improvement by imagining realistic future observations conditioned on ego-robot actions. While WMs can model distributions over futures, policy evalua…

Autonomous Driving

Curiosity-Driven Imagination: Discovering Plan Operators and Learning Associated Policies for Open-World Adaptation

2025-03-06 · Pierrick Lorang, Hong Lu, Matthias Scheutz

Adapting quickly to dynamic, uncertain environments-often called "open worlds"-remains a major challenge in robotics. Traditional Task and Motion Planning (TAMP) approaches struggle to cope with unforeseen changes, are d…

Motion PlanningTask and Motion Planning

Sora and V-JEPA Have Not Learned The Complete Real World Model -- A Philosophical Analysis of Video AIs Through the Theory of Productive Imagination

2024-05-06 · Jianqiu Zhang

Sora from Open AI has shown exceptional performance, yet it faces scrutiny over whether its technological prowess equates to an authentic comprehension of reality. Critics contend that it lacks a foundational grasp of th…

Philosophy