paper-with-me

Papers

Leveraging Structure from Motion to Learn Discriminative Codebooks for Scalable Landmark Classification

2013-06-01 · CVPR 2013 6 · Alessandro Bergamo, Sudipta N. Sinha, Lorenzo Torresani

In this paper we propose a new technique for learning a discriminative codebook for local feature descriptors, specifically designed for scalable landmark classification. The key contribution lies in exploiting the knowledge of correspondences within sets of feature descriptors during codebook learning. Feature correspondences are obtained using structure from motion (SfM) computation on Internet photo collections which serve as the training data. Our codebook is defined by a random forest that is trained to map corresponding feature descriptors into identical codes. Unlike prior forest-based codebook learning methods, we utilize fine-grained descriptor labels and address the challenge of training a forest with an extremely large number of labels. Our codebook is used with various existing feature encoding schemes and also a variant we propose for importanceweighted aggregation of local features. We evaluate our approach on a public dataset of 25 landmarks and our new dataset of 620 landmarks (614K images). Our approach significantly outperforms the state of the art in landmark classification. Furthermore, our method is memory efficient and scalable.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

General Classification

Similar Papers 제목 키워드 기반

Persistence Codebooks for Topological Data Analysis

2018-02-13 · Bartosz Zielinski, Michal Lipinski, Mateusz Juda, Matthias Zeppelzauer 외

Persistent homology (PH) is a rigorous mathematical theory that provides a robust descriptor of data in the form of persistence diagrams (PDs) which are 2D multisets of points. Their variable size makes them, however, di…

BIG-bench Machine LearningQuantizationTopological Data Analysis

Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generation

2024-12-01 · CVPR 2025 1 · Shuling Zhao, Fa-Ting Hong, Xiaoshui Huang, Dan Xu

Talking head video generation aims to generate a realistic talking head video that preserves the person's identity from a source image and the motion from a driving video. Despite the promising progress made in the field…

Video Generation

Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document Generation

2023-12-15 · Youngjoon Yoo, Jongwon Choi

This paper introduces a novel approach for topic modeling utilizing latent codebooks from Vector-Quantized Variational Auto-Encoder~(VQ-VAE), discretely encapsulating the rich information of the pre-trained embeddings su…

Image GenerationLanguage ModelingLanguage Modelling

VQDNA: Unleashing the Power of Vector Quantization for Multi-Species Genomic Sequence Modeling

2024-05-13 · Siyuan Li, Zedong Wang, Zicheng Liu, Di wu 외

Similar to natural language models, pre-trained genome language models are proposed to capture the underlying intricacies within genomes with unsupervised sequence modeling. They have become essential tools for researche…

Quantization

Discriminative Sparse Coding on Multi-Manifold for Data Representation and Classification

2012-08-19 · Jing-Yan Wang

Sparse coding has been popularly used as an effective data representation method in various applications, such as computer vision, medical imaging and bioinformatics, etc. However, the conventional sparse coding algorith…

General Classification