paper-with-me

Papers

Text Augmented Correlation Transformer For Few-shot Classification & Segmentation

2025-01-01 · CVPR 2025 1 · Srinivasa Rao Nandam, Sara Atito, ZhenHua Feng, Josef Kittler, Muhammad Awais

Foundation models like CLIP and ALIGN have transformed few-shot and zero-shot vision applications by fusing visual and textual data, yet the integrative few-shot classification and segmentation (FS-CS) task primarily leverages visual cues, overlooking the potential of textual support. In FS-CS scenarios, ambiguous object boundaries and overlapping classes often hinder model performance, as limited visual data struggles to fully capture high-level semantics. To bridge this gap, we present a novel multi-modal FS-CS framework that integrates textual cues into support data, facilitating enhanced semantic disambiguation and fine-grained segmentation. Our approach first investigates the unique contributions of exclusive text-based support, using only class labels to achieve FS-CS. This strategy alone achieves performance competitive with vision-only methods on FS-CS tasks, underscoring the power of textual cues in few-shot learning. Building on this, we introduce a dual-modal prediction mechanism that synthesizes insights from both textual and visual support sets, yielding robust multi-modal predictions. This integration significantly elevates FS-CS performance, with classification and segmentation improvements of +3.7/6.6% (1-way 1-shot) and +8.0/6.5% (2-way 1-shot) on COCO-20^i, and +2.2/3.8% (1-way 1-shot) and +4.3/4.0% (2-way 1-shot) on Pascal-5^i. Additionally, in weakly supervised FS-CS settings, our method surpasses visual-only benchmarks using textual support exclusively, further enhanced by our dual-modal predictions. By rethinking the role of text in FS-CS, our work establishes new benchmarks for multi-modal few-shot learning and demonstrates the efficacy of textual cues for improving model generalization and segmentation accuracy.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot Classification and SegmentationFew-Shot LearningSegmentation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

ClarET: Pre-training a Correlation-Aware Context-To-Event Transformer for Event-Centric Generation and Classification

2022-03-04 · ACL 2022 5 · Yucheng Zhou, Tao Shen, Xiubo Geng, Guodong Long 외

Generating new events given context with correlated ones plays a crucial role in many event-centric reasoning tasks. Existing works either limit their scope to specific scenarios or overlook event-level correlations. In …

counterfactualFew-Shot Learning

Zero-shot Image Captioning by Anchor-augmented Vision-Language Space Alignment

2022-11-14 · Junyang Wang, Yi Zhang, Ming Yan, Ji Zhang 외

CLIP (Contrastive Language-Image Pre-Training) has shown remarkable zero-shot transfer capabilities in cross-modal correlation tasks such as visual classification and image retrieval. However, its performance in cross-mo…

Computational EfficiencyImage CaptioningImage RetrievalRetrieval

Adaptive Multi-Scale Correlation Meta-Network for Few-Shot Remote Sensing Image Classification

2026-01-18 · Anurag Kaushish, Ayan Sar, Sampurna Roy, Sudeshna Chakraborty 외 arxiv

Few-shot learning in remote sensing remains challenging due to three factors: the scarcity of labeled data, substantial domain shifts, and the multi-scale nature of geospatial objects. To address these issues, we introdu…

Remote Sensing Image ClassificationFew-Shot Learning

Identifying Misinformation on YouTube through Transcript Contextual Analysis with Transformer Models

2023-07-22 · Christos Christodoulou, Nikos Salamanos, Pantelitsa Leonidou, Michail Papadakis 외

Misinformation on YouTube is a significant concern, necessitating robust detection strategies. In this paper, we introduce a novel methodology for video classification, focusing on the veracity of the content. We convert…

ArticlesClassificationFew-Shot LearningMisinformation+6

Don’t Miss the Labels: Label-semantic Augmented Meta-Learner for Few-Shot Text Classification

2021-08-01 · Findings (ACL) 2021 8 · Qiaoyang Luo, Lingqiao Liu, YuHao Lin, Wei zhang
Few-Shot Text Classificationtext-classificationText Classification