paper-with-me

홈 › Papers

DeVAn: Dense Video Annotation for Video-Language Models

2023-10-08 · Tingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fan, Ding Zhou, Huaibo Huang, Ran He, Hongxia Yang

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeVAn (Dense Video Annotation). The dataset contains 8.5K YouTube video clips of 20-60 seconds in duration and covers a wide range of topics and interests. Each video clip is independently annotated by 5 human annotators, producing both captions (1 sentence) and summaries (3-10 sentences). Given any video selected from the dataset and its corresponding ASR information, we evaluate visuallanguage models on either caption or summary generation that is grounded in both the visual and auditory content of the video. Additionally, models are also evaluated on caption- and summary-based retrieval tasks, where the summary-based retrieval task requires the identification of a target video given excerpts of a given summary. Given the novel nature of the paragraph-length video summarization task, we compared different existing evaluation metrics and their alignment with human preferences and found that model-based evaluation metrics provide more semantically-oriented and human-aligned evaluation. Finally, we benchmarked a wide range of current video-language models on DeVAn, and we aim for DeVAn to serve as a useful evaluation set in the age of large language models and complex multi-modal tasks. Code is available at https: //github.com/TK-21st/DeVAn.

📄 PDF Abstract BibTeX arXiv:2310.05060

Code (1)

tk-21st/devan 공식 구현 pytorch

Tasks

RetrievalSentenceVideo Summarization

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Attention Based Encoder Decoder Model for Video Captioning in Nepali (2023)

2023-12-12 · Kabita Parajuli, Shashidhar Ram Joshi

Video captioning in Nepali, a language written in the Devanagari script, presents a unique challenge due to the lack of existing academic work in this domain. This work develops a novel encoder-decoder paradigm for Nepal…

DecoderVideo CaptioningVideo Description

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation

2026-04-29 · Mingji Ge, Qirui Chen, Zeqian Li, Weidi Xie arxiv

Long-term video understanding requires interpreting complex temporal events and reasoning over procedural activities. While instructional video corpora, like HowTo100M, offer rich resources for model training, they prese…

Zero-shot GeneralizationDense Video CaptioningCross-Modal Retrieval

How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning

2026-05-16 · Bosung Kim, Ruiyi Wang, David Acuna, Jaehun Jung 외 arxiv

Scaling robot policy learning is bottlenecked by the cost of collecting demonstrations, while language annotations for existing demonstrations are comparatively cheap. We study language density as a lever for extracting …

Robot Manipulation

Multimodal Pretraining for Dense Video Captioning

2020-11-10 · Asian Chapter of the Association for Computational Linguistics 2020 · Gabriel Huang, Bo Pang, Zhenhai Zhu, Clara Rivera 외

Learning specific hands-on skills such as cooking, car maintenance, and home repairs increasingly happens via instructional videos. The user experience with such videos is known to be improved by meta-information such as…

Dense Video CaptioningVideo Captioning

Zero-Shot Dense Video Captioning by Jointly Optimizing Text and Moment

2023-07-05 · Yongrae Jo, Seongyun Lee, Aiden SJ Lee, Hyunji Lee 외

Dense video captioning, a task of localizing meaningful moments and generating relevant captions for videos, often requires a large, expensive corpus of annotated video segments paired with text. In an effort to minimize…

Dense Video CaptioningLanguage ModellingText GenerationVideo Captioning+1