paper-with-me

홈 › Papers

SketchQL Demonstration: Zero-shot Video Moment Querying with Sketches

2024-05-28 · Renzhi Wu, Pramod Chunduri, Dristi J Shah, Ashmitha Julius Aravind, Ali Payani, Xu Chu, Joy Arulraj, Kexin Rong

In this paper, we will present SketchQL, a video database management system (VDBMS) for retrieving video moments with a sketch-based query interface. This novel interface allows users to specify object trajectory events with simple mouse drag-and-drop operations. Users can use trajectories of single objects as building blocks to compose complex events. Using a pre-trained model that encodes trajectory similarity, SketchQL achieves zero-shot video moments retrieval by performing similarity searches over the video to identify clips that are the most similar to the visual query. In this demonstration, we introduce the graphic user interface of SketchQL and detail its functionalities and interaction mechanisms. We also demonstrate the end-to-end usage of SketchQL from query composition to video moments retrieval using real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2405.18334

Code (0)

등록된 구현이 없습니다.

Tasks

ManagementRetrieval

Similar Papers 제목 키워드 기반

Zero-Shot Dense Video Captioning by Jointly Optimizing Text and Moment

2023-07-05 · Yongrae Jo, Seongyun Lee, Aiden SJ Lee, Hyunji Lee 외

Dense video captioning, a task of localizing meaningful moments and generating relevant captions for videos, often requires a large, expensive corpus of annotated video segments paired with text. In an effort to minimize…

Dense Video CaptioningLanguage ModellingText GenerationVideo Captioning+1

Zero-shot Video Moment Retrieval With Off-the-Shelf Models

2022-11-03 · Anuj Diwan, Puyuan Peng, Raymond J. Mooney

For the majority of the machine learning community, the expensive nature of collecting high-quality human-annotated data and the inability to efficiently finetune very large state-of-the-art pretrained models on limited …

Moment RetrievalRetrieval

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models

2025-01-14 · Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Ming Li 외

The target of video moment retrieval (VMR) is predicting temporal spans within a video that semantically match a given linguistic query. Existing VMR methods based on multimodal large language models (MLLMs) overly rely …

Moment RetrievalRetrieval

ZS4IE: A toolkit for Zero-Shot Information Extraction with simple Verbalizations

2022-03-25 · NAACL (ACL) 2022 7 · Oscar Sainz, Haoling Qiu, Oier Lopez de Lacalle, Eneko Agirre 외

The current workflow for Information Extraction (IE) analysts involves the definition of the entities/relations of interest and a training corpus with annotated examples. In this demonstration we introduce a new workflow…

Natural Language InferenceZero-Shot Learning

World Action Models are Zero-shot Policies

2026-02-17 · Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng, Shenyuan Gao 외 arxiv

State-of-the-art Vision-Language-Action (VLA) models excel at semantic generalization but struggle to generalize to unseen physical motions in novel environments. We introduce DreamZero, a World Action Model (WAM) built …

Zero-shot Generalization