Contributions of Shape, Texture, and Color in Visual Recognition
We investigate the contributions of three important features of the human visual system (HVS)~ -- ~shape, texture, and color ~ -- ~to object classification. We build a humanoid vision engine (HVE) that explicitly and separately computes shape, texture, and color features from images. The resulting feature vectors are then concatenated to support the final classification. We show that HVE can summarize and rank-order the contributions of the three features to object recognition. We use human experiments to confirm that both HVE and humans predominantly use some specific features to support the classification of specific classes (e.g., texture is the dominant feature to distinguish a zebra from other quadrupeds, both for humans and HVE). With the help of HVE, given any environment (dataset), we can summarize the most important features for the whole task (task-specific; e.g., color is the most important feature overall for classification with the CUB dataset), and for each class (class-specific; e.g., shape is the most important feature to recognize boats in the iLab-20M dataset). To demonstrate more usefulness of HVE, we use it to simulate the open-world zero-shot learning ability of humans with no attribute labeling. Finally, we show that HVE can also simulate human imagination ability with the combination of different features. We will open-source the HVE engine and corresponding datasets.
Code (1)
Tasks
AttributeGeneral ClassificationObject RecognitionZero-Shot LearningSimilar Papers 제목 키워드 기반
Shape and Texture Recognition in Large Vision-Language Models
Shape and texture recognition is fundamental to visual perception. The ability to identify shapes regardless of orientation, texture, or context, and to recognize textures independently of their associated objects, is es…
3D Shape Recognition3D Shape RetrievalMaterial RecognitionTexture Image RetrievalCompressive Self-localization Using Relative Attribute Embedding
The use of relative attribute (e.g., beautiful, safe, convenient) -based image embeddings in visual place recognition, as a domain-adaptive compact image descriptor that is orthogonal to the typical approach of absolute …
AttributeVisual Place RecognitionSelf-supervised Visual Attribute Learning for Fashion Compatibility
Many self-supervised learning (SSL) methods have been successful in learning semantically meaningful visual representations by solving pretext tasks. However, prior work in SSL focuses on tasks like object recognition or…
AttributeObject RecognitionRetrievalSelf-Supervised Learning+1Texture features in medical image analysis: a survey
The texture is defined as spatial structure of the intensities of the pixels in an image that is repeated periodically in the whole image or regions, and makes the concept of the image. Texture, color and shape are three…
image-classificationImage ClassificationMedical Image AnalysisMedical Image Classification+2A 3D Morphable Model of Craniofacial Shape and Texture Variation
We present a fully automatic pipeline to train 3D Morphable Models (3DMMs), with contributions in pose normalisation, dense correspondence using both shape and texture information, and high quality, high resolution textu…
Optical Flow Estimation