paper-with-me

Papers

Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks

2024-09-27 · Min Yang, Zichen Zhang, LiMin Wang

With the development of video understanding, there is a proliferation of tasks for clip-level temporal video analysis, including temporal action detection (TAD), temporal action segmentation (TAS), and generic event boundary detection (GEBD). While task-specific video understanding models have exhibited outstanding performance in each task, there remains a dearth of a unified framework capable of simultaneously addressing multiple tasks, which is a promising direction for the next generation of AI. To this end, in this paper, we propose a single unified framework, coined as Temporal2Seq, to formulate the output of these temporal video understanding tasks as a sequence of discrete tokens. With this unified token representation, Temporal2Seq can train a generalist model within a single architecture on different video understanding tasks. In the absence of multi-task learning (MTL) benchmarks, we compile a comprehensive co-training dataset by borrowing the datasets from TAD, TAS, and GEBD tasks. We evaluate our Temporal2Seq generalist model on the corresponding test sets of three tasks, demonstrating that Temporal2Seq can produce reasonable results on various tasks and achieve advantages compared with single-task training on this framework. We also investigate the generalization performance of our generalist model on new datasets from different tasks, which yields superior performance to the specific model.

📄 PDF Abstract BibTeX arXiv:2409.18478

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionAction SegmentationBoundary DetectionGeneric Event Boundary DetectionMulti-Task LearningTemporal Action SegmentationVideo Understanding

Similar Papers 제목 키워드 기반

Bridging Video Understanding and Generation in a Unified Framework

2026-06-30 · Yuqi Wang, Runyi Li, Ruoyu Feng, Renjie Chen 외 arxiv

Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video domain remains largely underexplored. A central challenge is that video…

Video GenerationImage Generation

VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding

2026-01-12 · Jiapeng Shi, Junke Wang, Zuyao You, Bo He 외 arxiv

This paper presents VideoLoom, a unified Video Large Language Model (Video LLM) for joint spatial-temporal understanding. To facilitate the development of fine-grained spatial and temporal localization capabilities, we c…

Referring Video Object Segmentation

Video Understanding: From Geometry and Semantics to Unified Models

2026-03-18 · Zhaochong An, Zirui Li, Mingqiao Ye, Feng Qiao 외 arxiv

Video understanding aims to enable models to perceive, reason about, and interact with the dynamic visual world. In contrast to image understanding, video understanding inherently requires modeling temporal dynamics and …

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

2026-05-02 · Alejandro Aparcedo, Akash Kumar, Aaryan Garg, Dalton Pham 외 arxiv

Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity types, failing to capture the freeform, m…

STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding

2026-04-13 · Wenhao Li, Xueying Jiang, Gongjie Zhang, Xiaoqin Zhang 외 arxiv

4D point cloud videos capture rich spatial and temporal dynamics of scenes which possess unique values in various 4D understanding tasks. However, most existing methods work in the spatiotemporal domain where the underly…

Representation Learning3D Action RecognitionSemantic Segmentation