Learning Multi-Object Tracking and Segmentation from Automatic Annotations
In this work we contribute a novel pipeline to automatically generate training data, and to improve over state-of-the-art multi-object tracking and segmentation (MOTS) methods. Our proposed track mining algorithm turns raw street-level videos into high-fidelity MOTS training data, is scalable and overcomes the need of expensive and time-consuming manual annotation approaches. We leverage state-of-the-art instance segmentation results in combination with optical flow predictions, also trained on automatically harvested training data. Our second major contribution is MOTSNet - a deep learning, tracking-by-detection architecture for MOTS - deploying a novel mask-pooling layer for improved object association over time. Training MOTSNet with our automatically extracted data leads to significantly improved sMOTSA scores on the novel KITTI MOTS dataset (+1.9%/+7.5% on cars/pedestrians), and MOTSNet improves by +4.1% over previously best methods on the MOTSChallenge dataset. Our most impressive finding is that we can improve over previous best-performing works, even in complete absence of manually annotated MOTS training data.
Code (0)
등록된 구현이 없습니다.
Tasks
Instance SegmentationMulti-Object TrackingMulti-Object Tracking and SegmentationObjectObject TrackingOptical Flow EstimationSemantic SegmentationSimilar Papers 제목 키워드 기반
MOTS: Multi-Object Tracking and Segmentation
This paper extends the popular task of multi-object tracking to multi-object tracking and segmentation (MOTS). Towards this goal, we create dense pixel-level annotations for two existing tracking datasets using a semi-au…
Multi-Object TrackingMulti-Object Tracking and SegmentationMultiple Object TrackingObject+2Multi-Granularity Video Object Segmentation
Current benchmarks for video segmentation are limited to annotating only salient objects (i.e., foreground instances). Despite their impressive architectural designs, previous works trained on these benchmarks have strug…
ObjectSegmentationSemantic SegmentationVideo Object Segmentation+2Full segmentation annotations of 3D time-lapse microscopy images of MDA231 cells
High-quality, publicly available segmentation annotations of image and video datasets are critical for advancing the field of image processing. In particular, annotations of volumetric images of a large number of targets…
Cell SegmentationEntitySAM: Segment Everything in Video
Automatically tracking and segmenting every video entity remains a significant challenge. Despite rapid advancements in video segmentation, even state-of-the-art models like SAM 2 struggle to consistently track all e…
DecoderObjectSegmentationSemantic Segmentation+2Generating Masks from Boxes by Mining Spatio-Temporal Consistencies in Videos
Segmenting objects in videos is a fundamental computer vision task. The current deep learning based paradigm offers a powerful, but data-hungry solution. However, current datasets are limited by the cost and human effort…
ObjectSegmentationSemantic SegmentationVideo Object Segmentation+2