Towards efficient representation identification in supervised learning
Humans have a remarkable ability to disentangle complex sensory inputs (e.g., image, text) into simple factors of variation (e.g., shape, color) without much supervision. This ability has inspired many works that attempt to solve the following question: how do we invert the data generation process to extract those factors with minimal or no supervision? Several works in the literature on non-linear independent component analysis have established this negative result; without some knowledge of the data generation process or appropriate inductive biases, it is impossible to perform this inversion. In recent years, a lot of progress has been made on disentanglement under structural assumptions, e.g., when we have access to auxiliary information that makes the factors of variation conditionally independent. However, existing work requires a lot of auxiliary information, e.g., in supervised classification, it prescribes that the number of label classes should be at least equal to the total dimension of all factors of variation. In this work, we depart from these assumptions and ask: a) How can we get disentanglement when the auxiliary information does not provide conditional independence over the factors of variation? b) Can we reduce the amount of auxiliary information required for disentanglement? For a class of models where auxiliary information does not ensure conditional independence, we show theoretically and experimentally that disentanglement (to a large extent) is possible even when the auxiliary information dimension is much less than the dimension of the true latent representation.
Code (1)
Tasks
DisentanglementSimilar Papers 제목 키워드 기반
Decorrelation-based Self-Supervised Visual Representation Learning for Writer Identification
Self-supervised learning has developed rapidly over the last decade and has been applied in many areas of computer vision. Decorrelation-based self-supervised pretraining has shown great promise among non-contrastive alg…
Representation LearningSelf-Supervised LearningExploring Stronger Transformer Representation Learning for Occluded Person Re-Identification
Due to some complex factors (e.g., occlusion, pose variation and diverse camera perspectives), extracting stronger feature representation in person re-identification remains a challenging task. In this paper, we proposed…
Contrastive LearningOccluded Person Re-IdentificationPerson Re-IdentificationRepresentation LearningTransferring a Semantic Representation for Person Re-Identification and Search
Learning semantic attributes for person re-identification and description-based person search has gained increasing interest due to attributes' great potential as a pose and view-invariant representation. However, existi…
AttributePerson Re-IdentificationPerson SearchLabel Aware Speech Representation Learning For Language Identification
Speech representation learning approaches for non-semantic tasks such as language recognition have either explored supervised embedding extraction methods using a classifier model or self-supervised representation learni…
Language IdentificationMissing LabelsRepresentation Learningspeech-recognition+3Improved Language Identification Through Cross-Lingual Self-Supervised Learning
Language identification greatly impacts the success of downstream tasks such as automatic speech recognition. Recently, self-supervised speech representations learned by wav2vec 2.0 have been shown to be very effective f…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationSelf-Supervised Learning+2