paper-with-me

홈 › Papers

TubeTK: Adopting Tubes to Track Multi-Object in a One-Step Training Model

2020-06-10 · CVPR 2020 6 · Bo Pang, Yizhuo Li, Yifan Zhang, Muchen Li, Cewu Lu

Multi-object tracking is a fundamental vision problem that has been studied for a long time. As deep learning brings excellent performances to object detection algorithms, Tracking by Detection (TBD) has become the mainstream tracking framework. Despite the success of TBD, this two-step method is too complicated to train in an end-to-end manner and induces many challenges as well, such as insufficient exploration of video spatial-temporal information, vulnerability when facing object occlusion, and excessive reliance on detection results. To address these challenges, we propose a concise end-to-end model TubeTK which only needs one step training by introducing the ``bounding-tube" to indicate temporal-spatial locations of objects in a short video clip. TubeTK provides a novel direction of multi-object tracking, and we demonstrate its potential to solve the above challenges without bells and whistles. We analyze the performance of TubeTK on several MOT benchmarks and provide empirical evidence to show that TubeTK has the ability to overcome occlusions to some extent without any ancillary technologies like Re-ID. Compared with other methods that adopt private detection results, our one-stage end-to-end model achieves state-of-the-art performances even if it adopts no ready-made detection results. We hope that the proposed TubeTK model can serve as a simple but strong alternative for video-based MOT task. The code and models are available at https://github.com/BoPang1996/TubeTK.

📄 PDF Abstract BibTeX arXiv:2006.05683

Code (1)

BoPang1996/TubeTK 공식 구현 pytorch

Tasks

Multi-Object TrackingObjectobject-detectionObject DetectionObject Tracking

Similar Papers 제목 키워드 기반

A Low-Computational Video Synopsis Framework with a Standard Dataset

2024-09-08 · Ramtin Malekpour, M. Mehrdad Morsali, Hoda Mohammadzade

Video synopsis is an efficient method for condensing surveillance videos. This technique begins with the detection and tracking of objects, followed by the creation of object tubes. These tubes consist of sequences, each…

Objectobject-detectionObject DetectionObject Tracking+1

Deep-learning recognition and tracking of individual nanotubes in low-contrast microscopy videos

2024-10-17 · Vladimir Pimonov, Said Tahir, Vincent Jourdain

This study addresses the challenge of analyzing the growth kinetics of carbon nanotubes using in-situ homodyne polarization microscopy (HPM) by developing an automated deep learning (DL) approach. A Mask-RCNN architectur…

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation

2024-12-10 · Thong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T Nguyen 외

To equip artificial intelligence with a comprehensive understanding towards a temporal world, video and 4D panoptic scene graph generation abstracts visual data into nodes to represent entities and edges to capture tempo…

Contrastive LearningGraph GenerationPanoptic Scene Graph GenerationRelation+2

Object Detection, Tracking, and Motion Segmentation for Object-level Video Segmentation

2016-08-10 · Benjamin Drayer, Thomas Brox

We present an approach for object segmentation in videos that combines frame-level object detection with concepts from object tracking and motion segmentation. The approach extracts temporally consistent object tubes bas…

Motion SegmentationObjectobject-detectionObject Detection+5

Spatio-Temporal Action Detection with Multi-Object Interaction

2020-04-01 · Huijuan Xu, Lizhi Yang, Stan Sclaroff, Kate Saenko 외

Spatio-temporal action detection in videos requires localizing the action both spatially and temporally in the form of an "action tube". Nowadays, most spatio-temporal action detection datasets (e.g. UCF101-24, AVA, DALY…

Action DetectionHuman DetectionObjectregression