D2-Net: A Trainable CNN for Joint Description and Detection of Local Features
In this work we address the problem of finding reliable pixel-level correspondences under difficult imaging conditions. We propose an approach where a single convolutional neural network plays a dual role: It is simultaneously a dense feature descriptor and a feature detector. By postponing the detection to a later stage, the obtained keypoints are more stable than their traditional counterparts based on early detection of low-level structures. We show that this model can be trained using pixel correspondences extracted from readily available large-scale SfM reconstructions, without any further annotations. The proposed method obtains state-of-the-art performance on both the difficult Aachen Day-Night localization dataset and the InLoc indoor localization benchmark, as well as competitive performance on other benchmarks for image matching and 3D reconstruction.
Code (1)
Tasks
3D ReconstructionIndoor LocalizationSimilar Papers 제목 키워드 기반
D2-Net: A Trainable CNN for Joint Detection and Description of Local Features
In this work we address the problem of finding reliable pixel-level correspondences under difficult imaging conditions. We propose an approach where a single convolutional neural network plays a dual role: It is simultan…
3D ReconstructionImage MatchingIndoor LocalizationJoint Event Detection and Description in Continuous Video Streams
Dense video captioning is a fine-grained video understanding task that involves two sub-problems: localizing distinct events in a long video stream, and generating captions for the localized events. We propose the Joint …
Dense CaptioningDense Video CaptioningEvent DetectionVideo Captioning+1Multi-modal Retinal Image Registration Using a Keypoint-Based Vessel Structure Aligning Network
In ophthalmological imaging, multiple imaging systems, such as color fundus, infrared, fluorescein angiography, optical coherence tomography (OCT) or OCT angiography, are often involved to make a diagnosis of retinal dis…
Graph Neural NetworkImage RegistrationKeypoint DetectionWALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes
Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is absent, large vision-language models usually address this task. We ask …
Object LocalizationObject DetectionD3Feat: Joint Learning of Dense Detection and Description of 3D Local Features
A successful point cloud registration often lies on robust establishment of sparse matches through discriminative 3D local features. Despite the fast evolution of learning-based 3D feature descriptors, little attention h…
Point Cloud Registration