paper-with-me

홈 › Papers

From Red Wine to Red Tomato: Composition With Context

2017-07-01 · CVPR 2017 7 · Ishan Misra, Abhinav Gupta, Martial Hebert

Compositionality and contextuality are key building blocks of intelligence. They allow us to compose known concepts to generate new and complex ones. However, traditional learning methods do not model both these properties and require copious amounts of labeled data to learn new concepts. A large fraction of existing techniques, e.g., using late fusion, compose concepts but fail to model contextuality. For example, red in red wine is different from red in red tomatoes. In this paper, we present a simple method that respects contextuality in order to compose classifiers of known visual concepts. Our method builds upon the intuition that classifiers lie in a smooth space where compositional transforms can be modeled. We show how it can generalize to unseen combinations of concepts. Our results on composing attributes, objects as well as composing subject, predicate, and objects demonstrate its strong generalization performance compared to baselines. Finally, we present detailed analysis of our method and highlight its properties.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning

2025-02-01 · Zhixi Cai, Fucai Ke, Simindokht Jahangard, Maria Garcia de la Banda 외

Visual Grounding (VG) tasks, such as referring expression detection and segmentation tasks are important for linking visual entities to context, especially in complex reasoning tasks that require detailed query interpret…

Referring ExpressionVisual Grounding

Prompting Language-Informed Distribution for Compositional Zero-Shot Learning

2023-05-23 · Wentao Bao, Lichang Chen, Heng Huang, Yu Kong

Compositional zero-shot learning (CZSL) task aims to recognize unseen compositional visual concepts, e.g., sliced tomatoes, where the model is learned only from the seen compositions, e.g., sliced potatoes and red tomato…

Compositional Zero-Shot LearningInformativenessZero-shot GeneralizationZero-Shot Learning

Learning to Taste: A Multimodal Wine Dataset

2023-08-31 · NeurIPS 2023 11 · Thoranna Bender, Simon Moe Sørensen, Alireza Kashani, K. Eldjarn Hjorleifsson 외

We present WineSensed, a large multimodal wine dataset for studying the relations between visual perception, language, and flavor. The dataset encompasses 897k images of wine labels and 824k reviews of wines curated from…

MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning

2026-01-27 · Zhixi Cai, Fucai Ke, Kevin Leo, Sukai Huang 외 arxiv

Recent vision-language models have strong perceptual ability but their implicit reasoning is hard to explain and easily generates hallucinations on complex queries. Compositional methods improve interpretability, but mos…

Visual Reasoning

WiNet: Wavelet-based Incremental Learning for Efficient Medical Image Registration

2024-07-18 · Xinxing Cheng, Xi Jia, Wenqi Lu, Qiufu Li 외

Deep image registration has demonstrated exceptional accuracy and fast inference. Recent advances have adopted either multiple cascades or pyramid architectures to estimate dense deformation fields in a coarse-to-fine ma…

GPUImage RegistrationIncremental LearningMedical Image Registration