paper-with-me

홈 › Papers

Learning to Decompose Visual Features with Latent Textual Prompts

2022-10-09 · Feng Wang, Manling Li, Xudong Lin, Hairong Lv, Alexander G. Schwing, Heng Ji

Recent advances in pre-training vision-language models like CLIP have shown great potential in learning transferable visual representations. Nonetheless, for downstream inference, CLIP-like models suffer from either 1) degraded accuracy and robustness in the case of inaccurate text descriptions during retrieval-based inference (the challenge for zero-shot protocol); or 2) breaking the well-established vision-language alignment (the challenge for linear probing). To address them, we propose Decomposed Feature Prompting (DeFo). DeFo leverages a flexible number of learnable embeddings as textual input while maintaining the vision-language dual-model architecture, which enables the model to learn decomposed visual features with the help of feature-level textual prompts. We further use an additional linear layer to perform classification, allowing a scalable size of language inputs. Our empirical study shows DeFo's significance in improving the vision-language models. For example, DeFo obtains 73.2% test accuracy on ImageNet with a ResNet-50 backbone without tuning any pretrained weights of both the vision and language encoder, outperforming zero-shot CLIP by a large margin of 15.0%, and outperforming state-of-the-art vision-language prompt tuning method by 7.6%.

📄 PDF Abstract BibTeX arXiv:2210.04287

Code (0)

등록된 구현이 없습니다.

Tasks

Retrieval

Methods 이 논문이 사용한 방법론

Test 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.

Similar Papers 제목 키워드 기반

Inspiration Seeds: Learning Non-Literal Visual Combinations for Generative Exploration

2026-02-09 · Kfir Goldberg, Elad Richardson, Yael Vinker arxiv

While generative models have become powerful tools for image synthesis, they are typically optimized for executing carefully crafted textual prompts, offering limited support for the open-ended visual exploration that of…

Image Generation

Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs

2026-04-08 · Mengdan Zhu, Senhao Cheng, Liang Zhao arxiv

Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the cost of tool calls or rely on localized patch-based embeddings that a…

Visual Reasoning

RSRefSeg: Referring Remote Sensing Image Segmentation with Foundation Models

2025-01-12 · Keyan Chen, Jiafan Zhang, Chenyang Liu, Zhengxia Zou 외

Referring remote sensing image segmentation is crucial for achieving fine-grained visual understanding through free-format textual input, enabling enhanced scene and object extraction in remote sensing applications. Curr…

Image SegmentationSegmentationSemantic Segmentation

Latent Beam Diffusion Models for Decoding Image Sequences

2025-03-26 · Guilherme Fernandes, Vasco Ramos, Regev Cohen, Idan Szpektor 외

While diffusion models excel at generating high-quality images from text prompts, they struggle with visual consistency in image sequences. Existing methods generate each image independently, leading to disjointed narrat…

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models

2025-01-01 · Emily Johnson, Noah Wilson

Text-to-image generation has witnessed significant advancements with the integration of Large Vision-Language Models (LVLMs), yet challenges remain in aligning complex textual descriptions with high-quality, visually coh…

Image GenerationText to Image GenerationText-to-Image Generation