SWIFT: Sparse Withdrawal of Inliers in a First Trial
We study the simultaneous detection of multiple structures in the presence of overwhelming number of outliers in a large population of points. Our approach reduces the problem to sampling an extremely sparse subset of the original population of data in one grab, followed by an unsupervised clustering of the population based on a set of instantiated models from this sparse subset. We show that the problem can be modeled using a multivariate hypergeometric distribution, and derive accurate mathematical bounds to determine a tight approximation to the sample size, leading thus to a sparse sampling strategy. We evaluate the method thoroughly in terms of accuracy, its behavior against varying input parameters, and comparison against existing methods, including the state of the art. The key features of the proposed approach are: (i) sparseness of the sampled set, where the level of sparseness is independent of the population size and the distribution of data, (ii) robustness in the presence of overwhelming number of outliers, and (iii) unsupervised detection of all model instances, i.e. without requiring any prior knowledge of the number of embedded structures. To demonstrate the generic nature of the proposed method, we show experimental results on different computer vision problems, such as detection of physical structures e.g. lines, planes, etc., as well as more abstract structures such as fundamental matrices, and homographies in multi-body structure from motion.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringSimilar Papers 제목 키워드 기반
Sparse One-Time Grab Sampling of Inliers
Estimating structures in "big data" and clustering them are among the most fundamental problems in computer vision, pattern recognition, data mining, and many other other research fields. Over the past few decades, many …
ClusteringA Cautionary Tale on Integrating Studies with Disparate Outcome Measures for Causal Inference
Data integration approaches are increasingly used to enhance the efficiency and generalizability of studies. However, a key limitation of these methods is the assumption that outcome measures are identical across dataset…
Causal InferenceData IntegrationChanges in Retirement Savings During the COVID Pandemic
This paper documents changes in retirement saving patterns at the onset of the COVID-19 pandemic. We construct a large panel of U.S. tax data, including tens of millions of person-year observations, and measure retiremen…
MediSwift: Efficient Sparse Pre-trained Biomedical Language Models
Large language models (LLMs) are typically trained on general source data for various domains, but a recent surge in domain-specific LLMs has shown their potential to outperform general-purpose models in domain-specific …
Question AnsweringWithdrawal Success Estimation
Given a geometric Levy alpha-stable wealth process, a log-Levy alpha-stable lower bound is constructed for the terminal wealth of a regular investing schedule. Using a transformation, the lower bound is applied to a sche…