paper-with-me

Papers

CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging

2024-07-10 · Raza Imam, Mohammed Talha Alam, Umaima Rahman, Mohsen Guizani, Fakhri Karray

Existing vision-text contrastive learning models enhance representation transferability and support zero-shot prediction by matching paired image and caption embeddings while pushing unrelated pairs apart. However, astronomical image-label datasets are significantly smaller compared to general image and label datasets available from the internet. We introduce CosmoCLIP, an astronomical image-text contrastive learning framework precisely fine-tuned on the pre-trained CLIP model using SpaceNet and BLIP-based captions. SpaceNet, attained via FLARE, constitutes ~13k optimally distributed images, while BLIP acts as a rich knowledge extractor. The rich semantics derived from this SpaceNet and BLIP descriptions, when learned contrastively, enable CosmoCLIP to achieve superior generalization across various in-domain and out-of-domain tasks. Our results demonstrate that CosmoCLIP is a straightforward yet powerful framework, significantly outperforming CLIP in zero-shot classification and image-text retrieval tasks.

📄 PDF Abstract BibTeX arXiv:2407.07315

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningImage-text RetrievalRetrievalText Retrievalzero-shot-classificationZero-Shot Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
BLIP Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based…
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Applying Vision Transformers on Spectral Analysis of Astronomical Objects

2025-05-30 · Luis Felipe Strano Moraes, Ignacio Becker, Pavlos Protopapas, Guillermo Cabrera-Vives

We apply pre-trained Vision Transformers (ViTs), originally developed for image recognition, to the analysis of astronomical spectral data. By converting traditional one-dimensional spectra into two-dimensional image rep…

AstroVLM: Expert Multi-agent Collaborative Reasoning for Astronomical Imaging Quality Diagnosis

2026-04-17 · Yaohui Han, Tianshuo Wang, Zixi Zhao, Zhengchun Zhu 외 arxiv

Vision Language Models (VLMs) have been applied to several specific domains and have shown strong problem-solving capabilities. However, astronomical imaging, a quite complex problem involving multidisciplinary knowledge…

AstroLLaVA: towards the unification of astronomical data and natural language

2025-04-11 · Sharaf Zaman, Michael J. Smith, Pranav Khetarpal, Rishabh Chakrabarty 외

We present AstroLLaVA, a vision language model for astronomy that enables interaction with astronomical imagery through natural dialogue. By fine-tuning the LLaVA model on a diverse dataset of $\sim$30k images with capti…

AstronomyImage CaptioningLanguage ModelingLanguage Modelling+2

Effective Fine-Tuning of Vision-Language Models for Accurate Galaxy Morphology Analysis

2024-11-29 · Ruoqi Wang, Haitao Wang, Qiong Luo

Galaxy morphology analysis involves classifying galaxies by their shapes and structures. For this task, directly training domain-specific models on large, annotated astronomical datasets is effective but costly. In contr…

Contrastive Learning

At First Sight: Zero-Shot Classification of Astronomical Images with Large Multimodal Models

2024-06-24 · Dimitrios Tanoglidis, Bhuvnesh Jain

Vision-Language multimodal Models (VLMs) offer the possibility for zero-shot classification in astronomy: i.e. classification via natural language prompts, with no training. We investigate two models, GPT-4o and LLaVA-Ne…

AstronomyClassificationzero-shot-classificationZero-Shot Learning