ReMOTS: Self-Supervised Refining Multi-Object Tracking and Segmentation
We aim to improve the performance of Multiple Object Tracking and Segmentation (MOTS) by refinement. However, it remains challenging for refining MOTS results, which could be attributed to that appearance features are not adapted to target videos and it is also difficult to find proper thresholds to discriminate them. To tackle this issue, we propose a self-supervised refining MOTS (i.e., ReMOTS) framework. ReMOTS mainly takes four steps to refine MOTS results from the data association perspective. (1) Training the appearance encoder using predicted masks. (2) Associating observations across adjacent frames to form short-term tracklets. (3) Training the appearance encoder using short-term tracklets as reliable pseudo labels. (4) Merging short-term tracklets to long-term tracklets utilizing adopted appearance features and thresholds that are automatically obtained from statistical information. Using ReMOTS, we reached the $1^{st}$ place on CVPR 2020 MOTS Challenge 1, with an sMOTSA score of $69.9$.
Code (0)
등록된 구현이 없습니다.
Tasks
Multi-Object TrackingMulti-Object Tracking and SegmentationMultiple Object TrackingObjectObject TrackingSimilar Papers 제목 키워드 기반
SERPENT-VLM : Self-Refining Radiology Report Generation Using Vision Language Models
Radiology Report Generation (R2Gen) demonstrates how Multi-modal Large Language Models (MLLMs) can automate the creation of accurate and coherent radiological reports. Existing methods often hallucinate details in text-b…
Causal Language ModelingHallucinationLanguage ModelingLanguage ModellingUnsupervised Object Localization: Observing the Background to Discover Objects
Recent advances in self-supervised visual representation learning have paved the way for unsupervised methods tackling tasks such as object discovery and instance segmentation. However, discovering objects in an image wi…
Instance SegmentationObjectObject DiscoveryObject Localization+7Self-distilled Feature Aggregation for Self-supervised Monocular Depth Estimation
Self-supervised monocular depth estimation has received much attention recently in computer vision. Most of the existing works in literature aggregate multi-scale features for depth prediction via either straightforward …
Depth EstimationDepth PredictionMonocular Depth EstimationBilevel Data Curation for LLM Fine-tuning: Offline Selection and Online Self-Refining Generation
Supervised fine-tuning (SFT) datasets are critical to the downstream performance of large language models, yet they often contain low-quality or harmful question-response pairs. To improve SFT data quality, we develop a …
Bootstrapping Objectness from Videos by Relaxed Common Fate and Visual Grouping
We study learning object segmentation from unlabeled videos. Humans can easily segment moving objects without knowing what they are. The Gestalt law of common fate, i.e., what move at the same speed belong together, has …
Motion SegmentationObjectObject DiscoveryOptical Flow Estimation+6