Implicit Label Augmentation on Partially Annotated Clips via Temporally-Adaptive Features Learning
Partially annotated clips contain rich temporal contexts that can complement the sparse key frame annotations in providing supervision for model training. We present a novel paradigm called Temporally-Adaptive Features (TAF) learning that can utilize such data to learn better single frame models. By imposing distinct temporal change rate constraints on different factors in the model, TAF enables learning from unlabeled frames using context to enhance model accuracy. TAF generalizes "slow feature" learning and we present much stronger empirical evidence than prior works, showing convincing gains for the challenging semantic segmentation task over a variety of architecture designs and on two popular datasets. TAF can be interpreted as an implicit label augmentation method but is a more principled formulation compared to existing explicit augmentation techniques. Our work thus connects two promising methods that utilize partially annotated clips for single frame model training and can inspire future explorations in this direction.
Code (0)
등록된 구현이 없습니다.
Tasks
Semantic SegmentationSimilar Papers 제목 키워드 기반
Prior to Segment: Foreground Cues for Weakly Annotated Classes in Partially Supervised Instance Segmentation
Instance segmentation methods require large datasets with expensive and thus limited instance-level mask labels. Partially supervised instance segmentation aims to improve mask prediction with limited mask labels by util…
Instance SegmentationSegmentationSemantic SegmentationFree Performance Gain from Mixing Multiple Partially Labeled Samples in Multi-label Image Classification
Multi-label image classification datasets are often partially labeled where many labels are missing, posing a significant challenge to training accurate deep classifiers. However, the powerful Mixup sample-mixing data au…
BenchmarkingData Augmentationimage-classificationImage Classification+2HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization
This paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos. We refer to it as HACS (Human Action Clips and Segments). We leverage both consensus and…
Action ClassificationAction LocalizationAction RecognitionTemporal Action Localization+2DiffusionMTL: Learning Multi-Task Denoising Diffusion Model from Partially Annotated Data
Recently, there has been an increased interest in the practical problem of learning multiple dense scene understanding tasks from partially annotated data, where each training sample is only labeled for a subset of the t…
DenoisingScene UnderstandingKnowledge-Spreader: Learning Facial Action Unit Dynamics with Extremely Limited Labels
Recent studies on the automatic detection of facial action unit (AU) have extensively relied on large-sized annotations. However, manually AU labeling is difficult, time-consuming, and costly. Most existing semi-supervis…
Out-of-Distribution Generalization