Per-Sample Kernel Adaptation for Visual Recognition and Grouping
Object, action, or scene representations that are corrupted by noise significantly impair the performance of visual recognition. Typically, partial occlusion, clutter, or excessive articulation affects only a subset of all feature dimensions and, most importantly, different dimensions are corrupted in different samples. Nevertheless, the common approach to this problem in feature selection and kernel methods is to down-weight or eliminate entire training samples or the same dimensions of all samples. Thus, valuable signal is lost, resulting in suboptimal classification. Our goal is, therefore, to adjust the contribution of individual feature dimensions when comparing any two samples and computing their similarity. Consequently, per-sample selection of informative dimensions is directly integrated into kernel computation. The interrelated problems of learning the parameters of a kernel classifier and determining the informative components of each sample are then addressed in a joint objective function. The approach can be integrated into the learning stage of any kernel-based visual recognition problem and it does not affect the computational performance in the retrieval phase. Experiments on diverse challenges of action recognition in videos and indoor scene classification show the general applicability of the approach and its ability to improve learning of visual representations.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionAction Recognition In Videosfeature selectionGeneral ClassificationRetrievalScene ClassificationTemporal Action LocalizationSimilar Papers 제목 키워드 기반
A Novel Locally Linear KNN Model for Visual Recognition
This paper presents a novel locally linear KNN model with the goal of not only developing efficient representation and classification methods, but also establishing a relation between them so as to approximate some class…
Action RecognitionDensity EstimationFace RecognitionGeneral Classification+3Perceptual Group Tokenizer: Building Perception with Iterative Grouping
Human visual recognition system shows astonishing capability of compressing visual information into a set of tokens containing rich representations without label supervision. One critical driving principle behind it is p…
Representation LearningSelf-Supervised Image ClassificationSelf-Supervised LearningSuperpixelsDomain Adaptation with Soft-margin multiple feature-kernel learning beats Deep Learning for surveillance face recognition
Face recognition (FR) is the most preferred mode for biometric-based surveillance, due to its passive nature of detecting subjects, amongst all different types of biometric traits. FR under surveillance scenario does not…
Domain AdaptationFace RecognitionEvent Recognition in Videos by Learning from Heterogeneous Web Sources
In this work, we propose to leverage a large number of loosely labeled web videos (e.g., from YouTube) and web images (e.g., from Google/Bing image search) for visual event recognition in consumer videos without requirin…
Domain AdaptationImage RetrievalMulti-modal Egocentric Activity Recognition using Audio-Visual Features
Egocentric activity recognition in first-person videos has an increasing importance with a variety of applications such as lifelogging, summarization, assisted-living and activity tracking. Existing methods for this task…
Activity RecognitionEgocentric Activity RecognitionOptical Flow Estimation