Virtual to Real adaptation of Pedestrian Detectors
Pedestrian detection through Computer Vision is a building block for a multitude of applications. Recently, there was an increasing interest in Convolutional Neural Network-based architectures for the execution of such a task. One of these supervised networks' critical goals is to generalize the knowledge learned during the training phase to new scenarios with different characteristics. A suitably labeled dataset is essential to achieve this purpose. The main problem is that manually annotating a dataset usually requires a lot of human effort, and it is costly. To this end, we introduce ViPeD (Virtual Pedestrian Dataset), a new synthetically generated set of images collected with the highly photo-realistic graphical engine of the video game GTA V - Grand Theft Auto V, where annotations are automatically acquired. However, when training solely on the synthetic dataset, the model experiences a Synthetic2Real Domain Shift leading to a performance drop when applied to real-world images. To mitigate this gap, we propose two different Domain Adaptation techniques suitable for the pedestrian detection task, but possibly applicable to general object detection. Experiments show that the network trained with ViPeD can generalize over unseen real-world scenarios better than the detector trained over real-world data, exploiting the variety of our synthetic dataset. Furthermore, we demonstrate that with our Domain Adaptation techniques, we can reduce the Synthetic2Real Domain Shift, making closer the two domains and obtaining a performance improvement when testing the network over the real-world images. The code, the models, and the dataset are made freely available at https://ciampluca.github.io/viped/
Code (0)
등록된 구현이 없습니다.
Tasks
Domain Adaptationobject-detectionObject DetectionPedestrian DetectionSimilar Papers 제목 키워드 기반
MixedPeds: Pedestrian Detection in Unannotated Videos using Synthetically Generated Human-agents for Training
We present a new method for training pedestrian detectors on an unannotated set of images. We produce a mixed reality dataset that is composed of real-world background images and synthetically generated static human-agen…
Mixed RealityPedestrian DetectionScene-Specific Pedestrian Detection Based on Parallel Vision
As a special type of object detection, pedestrian detection in generic scenes has made a significant progress trained with large amounts of labeled training data manually. While the models trained with generic dataset wo…
object-detectionObject DetectionPedestrian DetectionLearning Scene-Specific Pedestrian Detectors Without Real Data
We consider the problem of designing a scene-specific pedestrian detector in a scenario where we have zero instances of real pedestrian data (i.e., no labeled real data or unsupervised real data). This scenario may arise…
Pedestrian DetectionMVUDA: Unsupervised Domain Adaptation for Multi-view Pedestrian Detection
We address multi-view pedestrian detection in a setting where labeled data is collected using a multi-camera setup different from the one used for testing. While recent multi-view pedestrian detectors perform well on the…
Domain AdaptationPedestrian DetectionUnsupervised Domain AdaptationContrastive-SDXL: Annotation-Preserving Night-Time Augmentation for Pedestrian Detection
Night-time pedestrian detection remains challenging because labelled night-time data are limited and large illumination differences make daytime-only trained detectors unreliable. Latent diffusion models (LDMs) provide a…
Image-to-Image TranslationSemantic correspondencePedestrian Detection