paper-with-me

Papers

TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models

2023-12-01 · Pengxiang Li, Kai Chen, Zhili Liu, Ruiyuan Gao, Lanqing Hong, Guo Zhou, Hua Yao, Dit-yan Yeung, Huchuan Lu, Xu Jia

Despite remarkable achievements in video synthesis, achieving granular control over complex dynamics, such as nuanced movement among multiple interacting objects, still presents a significant hurdle for dynamic world modeling, compounded by the necessity to manage appearance and disappearance, drastic scale changes, and ensure consistency for instances across frames. These challenges hinder the development of video generation that can faithfully mimic real-world complexity, limiting utility for applications requiring high-level realism and controllability, including advanced scene simulation and training of perception systems. To address that, we propose TrackDiffusion, a novel video generation framework affording fine-grained trajectory-conditioned motion control via diffusion models, which facilitates the precise manipulation of the object trajectories and interactions, overcoming the prevalent limitation of scale and continuity disruptions. A pivotal component of TrackDiffusion is the instance enhancer, which explicitly ensures inter-frame consistency of multiple objects, a critical factor overlooked in the current literature. Moreover, we demonstrate that generated video sequences by our TrackDiffusion can be used as training data for visual perception models. To the best of our knowledge, this is the first work to apply video diffusion models with tracklet conditions and demonstrate that generated frames can be beneficial for improving the performance of object trackers.

📄 PDF Abstract BibTeX arXiv:2312.00651

Code (1)

pixeli99/svd_xtend pytorch

Tasks

Image ClassificationMulti-Object TrackingObjectobject-detectionObject DetectionObject TrackingVideo Generation

Methods 이 논문이 사용한 방법론

Copy-Paste 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Integrated Object Detection and Tracking with Tracklet-Conditioned Detection

2018-11-27 · Zheng Zhang, Dazhi Cheng, Xizhou Zhu, Stephen Lin 외

Accurate detection and tracking of objects is vital for effective video understanding. In previous work, the two tasks have been combined in a way that tracking is based heavily on detection, but the detection benefits m…

Objectobject-detectionObject DetectionVideo Object Detection+1

OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding

2025-04-15 · Dianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu 외

In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff, aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats al…

Semantic SegmentationVideo GenerationVideo Understanding

Video Relation Detection via Tracklet based Visual Transformer

2021-08-19 · Kaifeng Gao, Long Chen, Yifeng Huang, Jun Xiao

Video Visual Relation Detection (VidVRD), has received significant attention of our community over recent years. In this paper, we apply the state-of-the-art video object tracklet detection pipeline MEGA and deepSORT to …

DecoderRelationVideo Visual Relation Detection

Unleashing the Potential of Tracklets for Unsupervised Video Person Re-Identification

2024-06-20 · Nanxing Meng, Qizao Wang, Bin Li, xiangyang xue

With rich temporal-spatial information, video-based person re-identification methods have shown broad prospects. Although tracklets can be easily obtained with ready-made tracking models, annotating identities is still e…

Person Re-IdentificationVideo-Based Person Re-Identification

Temporally Coherent Bayesian Models for Entity Discovery in Videos by Tracklet Clustering

2014-09-22 · Adway Mitra, Soma Biswas, Chiranjib Bhattacharyya

A video can be represented as a sequence of tracklets, each spanning 10-20 frames, and associated with one entity (eg. a person). The task of \emph{Entity Discovery} in videos can be naturally posed as tracklet clusterin…

ClusteringVideo Summarization