LOCL: Learning Object-Attribute Composition using Localization
This paper describes LOCL (Learning Object Attribute Composition using Localization) that generalizes composition zero shot learning to objects in cluttered and more realistic settings. The problem of unseen Object Attribute (OA) associations has been well studied in the field, however, the performance of existing methods is limited in challenging scenes. In this context, our key contribution is a modular approach to localizing objects and attributes of interest in a weakly supervised context that generalizes robustly to unseen configurations. Localization coupled with a composition classifier significantly outperforms state of the art (SOTA) methods, with an improvement of about 12% on currently available challenging datasets. Further, the modularity enables the use of localized feature extractor to be used with existing OA compositional learning methods to improve their overall performance.
Code (1)
Tasks
AttributeObjectZero-Shot LearningSimilar Papers 제목 키워드 기반
LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model
The capacity of existing human keypoint localization models is limited by keypoint priors provided by the training data. To alleviate this restriction and pursue more general model, this work studies keypoint localizatio…
Language ModelingLanguage ModellingLarge Language ModelLeveraging Large Language Models for Effective Label-free Node Classification in Text-Attributed Graphs
Graph neural networks (GNNs) have become the preferred models for node classification in graph data due to their robust capabilities in integrating graph structures and attributes. However, these models heavily depend on…
ClassificationNode ClassificationPoseLLM: Enhancing Language-Guided Human Pose Estimation with MLP Alignment
Human pose estimation traditionally relies on architectures that encode keypoint priors, limiting their generalization to novel poses or unseen keypoints. Recent language-guided approaches like LocLLM reformulate keypoin…
Large Language ModelPose EstimationZero-shot Generalizationnextlocllm: next location prediction using LLMs
Next location prediction is a critical task in human mobility analysis and serves as a foundation for various downstream applications. Existing methods typically rely on discrete IDs to represent locations, which inheren…
PredictionAttributed Grammars for Joint Estimation of Human Attributes, Part and Pose
In this paper, we are interested in developing compositional models to explicit representing pose, parts and attributes and tackling the tasks of attribute recognition, pose estimation and part localization jointly. Thi…
AttributeHuman ParsingPose Estimation