Papers zero-shot-classification
“zero-shot-classification” 태그가 달린 논문 422편 · 필터 해제
DEARLi: Decoupled Enhancement of Recognition and Localization for Semi-supervised Panoptic Segmentation
Pixel-level annotation is expensive and time-consuming. Semi-supervised segmentation methods address this challenge by learning models on few labeled images alongside a large corpus of unlabeled images. Although foundati…
DecoderGPUPanoptic SegmentationSemantic Segmentation+3Comparison of ConvNeXt and Vision-Language Models for Breast Density Assessment in Screening Mammography
Mammographic breast density classification is essential for cancer risk assessment but remains challenging due to subjective interpretation and inter-observer variability. This study compares multimodal and CNN-based met…
breast density classificationClassificationzero-shot-classificationZero-Shot LearningHarmonizing and Merging Source Models for CLIP-based Domain Generalization
CLIP-based domain generalization aims to improve model generalization to unseen domains by leveraging the powerful zero-shot classification capabilities of CLIP and multiple source datasets. Existing methods typically tr…
Domain Generalizationzero-shot-classificationZero-Shot LearningEfficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
Medical vision-language alignment through cross-modal contrastive learning shows promising performance in image-text matching tasks, such as retrieval and zero-shot classification. However, conventional cross-modal contr…
Contrastive LearningImage-text matchingImage to textImage-to-Text Retrieval+5GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models
Classifying geospatial imagery remains a major bottleneck for applications such as disaster response and land-use monitoring-particularly in regions where annotated data is scarce or unavailable. Existing tools (e.g., RS…
ClassificationDisaster Responseimage-classificationImage Classification+6AmorLIP: Efficient Language-Image Pretraining via Amortization
Contrastive Language-Image Pretraining (CLIP) has demonstrated strong zero-shot performance across diverse downstream text-image tasks. Existing CLIP methods typically optimize a contrastive objective using negative samp…
Contrastive LearningRepresentation Learningzero-shot-classificationZero-Shot LearningDistill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
We present Distill CLIP (DCLIP), a fine-tuned variant of the CLIP model that enhances multimodal image-text retrieval while preserving the original model's strong zero-shot classification capabilities. CLIP models are ty…
Contrastive LearningImage-text RetrievalRetrievalText Retrieval+2Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment
We introduce Perceptual-Initialization (PI), a paradigm shift in visual representation learning that incorporates human perceptual structure during the initialization phase rather than as a downstream fine-tuning step. B…
Representation LearningRetrievalSelf-Supervised LearningTriplet+2Uniformity First: Uniformity-aware Test-time Adaptation of Vision-language Models against Image Corruption
Pre-trained vision-language models such as contrastive language-image pre-training (CLIP) have demonstrated a remarkable generalizability, which has enabled a wide range of applications represented by zero-shot classific…
Knowledge DistillationTest-time Adaptationzero-shot-classificationZero-Shot LearningStarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment
Learning robust representations from data often requires scale, which has led to the success of recent zero-shot models such as CLIP. However, the obtained robustness can easily be deteriorated when these models are fine…
zero-shot-classificationZero-Shot LearningFrom Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
Pretrained vision-language models (VLMs), e.g., CLIP, demonstrate impressive zero-shot capabilities on downstream tasks. Prior research highlights the crucial role of visual augmentation techniques, like random cropping,…
feature selectionOut-of-Distribution GeneralizationTest-time Adaptationzero-shot-classification+1Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner
Recent advances in vision language models (VLMs) have enabled broad progress in the general medical field. However, pathology still remains a more challenging subdomain, with current pathology specific VLMs exhibiting li…
Cross-Modal RetrievalDiagnosticImage DescriptionMultimodal Reasoning+5Advanced Crash Causation Analysis for Freeway Safety: A Large Language Model Approach to Identifying Key Contributing Factors
Understanding the factors contributing to traffic crashes and developing strategies to mitigate their severity is essential. Traditional statistical methods and machine learning models often struggle to capture the compl…
Language ModelingLanguage ModellingLarge Language Modelzero-shot-classification+1TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
Learning to associate audio with textual descriptions is valuable for a range of tasks, including pretraining, zero-shot classification, audio retrieval, audio captioning, and text-conditioned audio generation. Existing …
Audio captioningAudio GenerationSentencezero-shot-classification+1Image Classification Using a Diffusion Model as a Pre-Training Model
In this paper, we propose a diffusion model that integrates a representation-conditioning mechanism, where the representations derived from a Vision Transformer (ViT) are used to condition the internal process of a Trans…
Contrastive Learningimage-classificationImage Classificationmodel+3MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks
Medical vision-language models (VLMs) have shown promise as clinical assistants across various medical fields. However, specialized dermatology VLM capable of delivering professional and detailed diagnostic analysis rema…
DiagnosticInstruction FollowingLanguage ModelingLanguage Modelling+4FG-CLIP: Fine-Grained Visual and Textual Alignment
Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short c…
Image-text Retrievalobject-detectionObject DetectionOpen-vocabulary object detection+5Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning
Vision-language models (VLMs) allow to embed texts and images in a shared representation space. However, it has been shown that these models are subject to a modality gap phenomenon meaning there exists a clear separatio…
Representation Learningzero-shot-classificationZero-Shot LearningHelping Large Language Models Protect Themselves: An Enhanced Filtering and Summarization System
The recent growth in the use of Large Language Models has made them vulnerable to sophisticated adversarial assaults, manipulative prompts, and encoded malicious inputs. Existing countermeasures frequently necessitate re…
zero-shot-classificationZero-Shot LearningOn the effectiveness of Large Language Models in the mechanical design domain
In this work, we seek to understand the performance of large language models in the mechanical engineering domain. We leverage the semantic data found in the ABC dataset, specifically the assembly names that designers as…
ClassificationSentenceSentence-Pair Classificationzero-shot-classification+1