Temporal Coherence for Active Learning in Videos
Autonomous driving systems require huge amounts of data to train. Manual annotation of this data is time-consuming and prohibitively expensive since it involves human resources. Therefore, active learning emerged as an alternative to ease this effort and to make data annotation more manageable. In this paper, we introduce a novel active learning approach for object detection in videos by exploiting temporal coherence. Our active learning criterion is based on the estimated number of errors in terms of false positives and false negatives. The detections obtained by the object detector are used to define the nodes of a graph and tracked forward and backward to temporally link the nodes. Minimizing an energy function defined on this graphical model provides estimates of both false positives and false negatives. Additionally, we introduce a synthetic video dataset, called SYNTHIA-AL, specially designed to evaluate active learning for video object detection in road scenes. Finally, we show that our approach outperforms active learning baselines tested on two datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Active LearningAutonomous DrivingObjectobject-detectionObject DetectionVideo Object DetectionSimilar Papers 제목 키워드 기반
EmoCo: Visual Analysis of Emotion Coherence in Presentation Videos
Emotions play a key role in human communication and public presentations. Human emotions are usually expressed through multiple modalities. Therefore, exploring multimodal emotions and their coherence is of great value f…
ClusteringSentenceVideo Face Editing Using Temporal-Spatial-Smooth Warping
Editing faces in videos is a popular yet challenging aspect of computer vision and graphics, which encompasses several applications including facial attractiveness enhancement, makeup transfer, face replacement, and expr…
Exploring Temporal Coherence for More General Video Face Forgery Detection
Although current face manipulation techniques achieve impressive performance regarding quality and controllability, they are struggling to generate temporal coherent face videos. In this work, we explore to take full adv…
DeepFake DetectionEdit Temporal-Consistent Videos with Image Diffusion Model
Large-scale text-to-image (T2I) diffusion models have been extended for text-guided video editing, yielding impressive zero-shot video editing performance. Nonetheless, the generated videos usually show spatial irregular…
modelVideo EditingVideo Temporal ConsistencyEAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence
Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insuff…
Graph structure learningComputational EfficiencyTalking Head GenerationVideo Generation