Decorrelating Semantic Visual Attributes by Resisting the Urge to Share
Existing methods to learn visual attributes are prone to learning the wrong thing---namely, properties that are correlated with the attribute of interest among training samples. Yet, many proposed applications of attributes rely on being able to learn the correct semantic concept corresponding to each attribute. We propose to resolve such confusions by jointly learning decorrelated, discriminative attribute models. Leveraging side information about semantic relatedness, we develop a multi-task learning approach that uses structured sparsity to encourage feature competition among unrelated attributes and feature sharing among related attributes. On three challenging datasets, we show that accounting for structure in the visual attribute space is key to learning attribute models that preserve semantics, yielding improved generalizability that helps in the recognition and discovery of unseen object categories.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeMulti-Task LearningSimilar Papers 제목 키워드 기반
A Knowledge-guided Adversarial Defense for Resisting Malicious Visual Manipulation
Malicious applications of visual manipulation have raised serious threats to the security and reputation of users in many fields. To alleviate these issues, adversarial noise-based defenses have been enthusiastically stu…
Adversarial DefenseSecure Video Quality Assessment Resisting Adversarial Attacks
The exponential surge in video traffic has intensified the imperative for Video Quality Assessment (VQA). Leveraging cutting-edge architectures, current VQA models have achieved human-comparable accuracy. However, recent…
Adversarial DefenseVideo Quality AssessmentVisual Question Answering (VQA)Decorrelation using Optimal Transport
Being able to decorrelate a feature space from protected attributes is an area of active research and study in ethics, fairness, and also natural sciences. We introduce a novel decorrelation method using Convex Neural Op…
Binary ClassificationEthicsFairnessNormalising FlowsCausal Evidence for Attention Head Imbalance in Modality Conflict Hallucination
Modality-conflict hallucination occurs when multimodal large language models (MLLMs) prioritize erroneous textual premises over contradictory visual evidence. To understand why visual evidence fails to prevail during gen…
Dual Relation Mining Network for Zero-Shot Learning
Zero-shot learning (ZSL) aims to recognize novel classes through transferring shared semantic knowledge (e.g., attributes) from seen classes to unseen classes. Recently, attention-based methods have exhibited significant…
AttributeRelationTransfer LearningZero-Shot Learning