Integrative Few-Shot Learning for Classification and Segmentation
We introduce the integrative task of few-shot classification and segmentation (FS-CS) that aims to both classify and segment target objects in a query image when the target classes are given with a few examples. This task combines two conventional few-shot learning problems, few-shot classification and segmentation. FS-CS generalizes them to more realistic episodes with arbitrary image pairs, where each target class may or may not be present in the query. To address the task, we propose the integrative few-shot learning (iFSL) framework for FS-CS, which trains a learner to construct class-wise foreground maps for multi-label classification and pixel-wise segmentation. We also develop an effective iFSL model, attentive squeeze network (ASNet), that leverages deep semantic correlation and global self-attention to produce reliable foreground maps. In experiments, the proposed method shows promising performance on the FS-CS task and also achieves the state of the art on standard few-shot segmentation benchmarks.
Code (1)
Tasks
ClassificationFew-Shot Classification and SegmentationFew-Shot LearningFew-Shot Semantic SegmentationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONSegmentationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Text Augmented Correlation Transformer For Few-shot Classification & Segmentation
Foundation models like CLIP and ALIGN have transformed few-shot and zero-shot vision applications by fusing visual and textual data, yet the integrative few-shot classification and segmentation (FS-CS) task primarily…
Few-Shot Classification and SegmentationFew-Shot LearningSegmentationIGCN: Integrative Graph Convolution Networks for patient level insights and biomarker discovery in multi-omics integration
Developing computational tools for integrative analysis across multiple types of omics data has been of immense importance in cancer molecular biology and precision medicine research. While recent advancements have yield…
Node ClassificationAM-RADIO: Agglomerative Vision Foundation Model -- Reduce All Domains Into One
A handful of visual foundation models (VFMs) have recently emerged as the backbones for numerous downstream tasks. VFMs like CLIP, DINOv2, SAM are trained with distinct objectives, exhibiting unique characteristics for v…
AllBenchmarkingobject-detectionObject Detection+1AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into One
A handful of visual foundation models (VFMs) have recently emerged as the backbones for numerous downstream tasks. VFMs like CLIP DINOv2 SAM are trained with distinct objectives exhibiting unique characteristics for …
AllBenchmarkingobject-detectionObject Detection+1InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding
Grounding large language models (LLMs) in external knowledge sources is a promising method for faithful prediction. While existing grounding approaches work well for simple queries, many real-world information needs requ…