paper-with-me

홈 › Papers

TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

2024-11-27 · Shimin Chen, Xiaohan Lan, Yitian Yuan, Zequn Jie, Lin Ma

Rapid development of large language models (LLMs) has significantly advanced multimodal large language models (LMMs), particularly in vision-language tasks. However, existing video-language models often overlook precise temporal localization and struggle with videos of varying lengths. We introduce TimeMarker, a versatile Video-LLM designed for high-quality dialogue based on video content, emphasizing temporal localization. TimeMarker integrates Temporal Separator Tokens to enhance temporal awareness, accurately marking specific moments within videos. It employs the AnyLength mechanism for dynamic frame sampling and adaptive token merging, enabling effective handling of both short and long videos. Additionally, TimeMarker utilizes diverse datasets, including further transformed temporal-related video QA datasets, to bolster its temporal understanding capabilities. Image and interleaved data are also employed to further enhance the model's semantic perception ability. Evaluations demonstrate that TimeMarker achieves state-of-the-art performance across multiple benchmarks, excelling in both short and long video categories. Our project page is at \url{https://github.com/TimeMarker-LLM/TimeMarker/}.

📄 PDF Abstract BibTeX arXiv:2411.18211

Code (1)

timemarker-llm/timemarker 공식 구현

Tasks

Temporal LocalizationVideo Understanding

Similar Papers 제목 키워드 기반

Tube-Link: A Flexible Cross Tube Framework for Universal Video Segmentation

2023-03-22 · ICCV 2023 1 · Xiangtai Li, Haobo Yuan, Wenwei Zhang, Guangliang Cheng 외

Video segmentation aims to segment and track every pixel in diverse scenarios accurately. In this paper, we present Tube-Link, a versatile framework that addresses multiple core tasks of video segmentation with a unified…

Contrastive LearningSegmentationVideo Instance SegmentationVideo Panoptic Segmentation+2

MoVQA: A Benchmark of Versatile Question-Answering for Long-Form Movie Understanding

2023-12-08 · Hongjie Zhang, Yi Liu, Lu Dong, Yifei HUANG 외

While several long-form VideoQA datasets have been introduced, the length of both videos used to curate questions and sub-clips of clues leveraged to answer those questions have not yet reached the criteria for genuine l…

FormQuestion AnsweringVideo Question AnsweringVideo Understanding

Coding Standards as Anchors for the CVPR CLIC video track

2021-05-20 · Théo Ladune, Pierrick Philippe

In 2021, a new track has been initiated in the Challenge for Learned Image Compression~: the video track. This category proposes to explore technologies for the compression of short video clips at 1 Mbit/s. This paper pr…

Image Compression

Lotus: Creating Short Videos From Long Videos With Abstractive and Extractive Summarization

2025-02-10 · Aadit Barua, Karim Benharrak, Meng Chen, Mina Huh 외

Short-form videos are popular on platforms like TikTok and Instagram as they quickly capture viewers' attention. Many creators repurpose their long-form videos to produce short-form videos, but creators report that plann…

Extractive SummarizationForm

Turbo Training with Token Dropout

2022-10-10 · Tengda Han, Weidi Xie, Andrew Zisserman

The objective of this paper is an efficient training method for video tasks. We make three contributions: (1) We propose Turbo training, a simple and versatile training paradigm for Transformers on multiple video tasks. …

Action ClassificationClassificationRepresentation Learning