paper-with-me

홈 › Papers

Reasoning about Fine-grained Attribute Phrases using Reference Games

2017-08-29 · ICCV 2017 10 · Jong-Chyi Su, Chenyun Wu, Huaizu Jiang, Subhransu Maji

We present a framework for learning to describe fine-grained visual differences between instances using attribute phrases. Attribute phrases capture distinguishing aspects of an object (e.g., "propeller on the nose" or "door near the wing" for airplanes) in a compositional manner. Instances within a category can be described by a set of these phrases and collectively they span the space of semantic attributes for a category. We collect a large dataset of such phrases by asking annotators to describe several visual differences between a pair of instances within a category. We then learn to describe and ground these phrases to images in the context of a *reference game* between a speaker and a listener. The goal of a speaker is to describe attributes of an image that allows the listener to correctly identify it within a pair. Data collected in a pairwise manner improves the ability of the speaker to generate, and the ability of the listener to interpret visual descriptions. Moreover, due to the compositionality of attribute phrases, the trained listeners can interpret descriptions not seen during training for image retrieval, and the speakers can generate attribute-based explanations for differences between previously unseen categories. We also show that embedding an image into the semantic space of attribute phrases derived from listeners offers 20% improvement in accuracy over existing attribute-based representations on the FGVC-aircraft dataset.

📄 PDF Abstract BibTeX arXiv:1708.08874

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeImage RetrievalRetrieval

Similar Papers 제목 키워드 기반

Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models

2023-08-18 · Navid Rajabi, Jana Kosecka

Large vision-and-language models (VLMs) trained to match images with text on large-scale datasets of image-text pairs have shown impressive generalization ability on several vision and language tasks. Several recent work…

Image-text matchingObject LocalizationQuestion AnsweringSpatial Reasoning+3

Toward Explainable and Fine-Grained 3D Grounding through Referring Textual Phrases

2022-07-05 · Zhihao Yuan, Xu Yan, Zhuo Li, Xuhao Li 외

Recent progress in 3D scene understanding has explored visual grounding (3DVG) to localize a target object through a language description. However, existing methods only consider the dependency between the entire sentenc…

ObjectRepresentation LearningScene UnderstandingSentence+1

Democratizing Fine-grained Visual Recognition with Large Language Models

2024-01-24 · Mingxuan Liu, Subhankar Roy, Wenjing Li, Zhun Zhong 외

Identifying subordinate-level categories from images is a longstanding task in computer vision and is referred to as fine-grained visual recognition (FGVR). It has tremendous significance in real-world applications since…

Fine-Grained Visual RecognitionWorld Knowledge

SCARE ― The Sentiment Corpus of App Reviews with Fine-grained Annotations in German

2016-05-01 · LREC 2016 5 · Mario S{\"a}nger, Ulf Leser, Steffen Kemmerer, Peter Adolphs 외

The automatic analysis of texts containing opinions of users about, e.g., products or political views has gained attention within the last decades. However, previous work on the task of analyzing user reviews about mobil…

Sentiment Analysis

MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence

2025-05-15 · Chonghan Liu, Haoran Wang, Felix Henry, Pu Miao 외

Spatial perception and reasoning are core components of human cognition, encompassing object recognition, spatial relational understanding, and dynamic reasoning. Despite progress in computer vision, existing benchmarks …

AttributeObjectObject RecognitionRelation+1