paper-with-me

홈 › Papers

CropVLM: A Domain-Adapted Vision-Language Model for Open-Set Crop Analysis

2026-05-05 · Abderrahmene Boudiaf, Sajd Javed arxiv

High-throughput plant phenotyping, the quantitative measurement of observable plant traits, is critical for modern breeding but remains constrained by a "phenotyping bottleneck," where manual data collection is labor-intensive and prone to observer bias. Conventional closed-set computer vision systems fail to address this challenge, as they require extensive species-specific annotation and lack the flexibility to handle diverse breeding populations. To bridge this gap, we present CropVLM, a Vision-Language Model (VLM) adapted for the agricultural domain via Domain-Specific Semantic Alignment (DSSA). Trained on 52,987 manually selected image-caption pairs covering 37 species in natural field conditions, CropVLM effectively maps agronomic terminology to fine-grained visual features. We further introduce the Hybrid Open-Set Localization Network (HOS-Net), an architecture that integrates CropVLM to enable the detection of novel crops solely from natural language descriptions without retraining. By eliminating the reliance on species-specific training data, CropVLM provides a scalable solution for high-throughput phenotyping, accelerating genetic gain and facilitating large-scale biodiversity research essential for sustainable agriculture. The trained model weights and complete pipeline implementation are publicly available at: https://github.com/boudiafA/CropVLM. In comprehensive evaluations, CropVLM achieves 72.51% zero-shot classification accuracy, outperforming seven CLIP-style baselines. Our detection pipeline demonstrates superior zero-shot generalization to novel species, achieving 49.17 AP50 on our CVTCropDet benchmark and 50.73 AP50 on tropical fruit species, compared to 34.89 and 48.58 for the next-best method, respectively.

📄 PDF Abstract BibTeX arXiv:2605.03259

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot Generalization

Similar Papers 제목 키워드 기반

CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception

2025-11-25 · Miguel Carvalho, Helder Dias, Bruno Martins arxiv

Vision-Language Models (VLMs) often struggle with tasks that require fine-grained image understanding, such as scene-text recognition or document analysis, due to perception limitations and visual fragmentation. To addre…

Reinforcement Learning

Domain Adapted Large Language Models for Additive Manufacturing

2026-03-23 · Peter Pak, Amir Barati Farimani arxiv

This work presents a collection of multi-modal domain adapted large language models built upon the instruction tuned variants of open weight models (Gemma 3, Qwen 3, Gemma 4) using a relatively small dataset of around 50…

PIR: Remote Sensing Image-Text Retrieval with Prior Instruction Representation Learning

2024-05-16 · Jiancheng Pan, Muyuan Ma, Qing Ma, Cong Bai 외

Remote sensing image-text retrieval constitutes a foundational aspect of remote sensing interpretation tasks, facilitating the alignment of vision and language representations. This paper introduces a prior instruction r…

Image-text RetrievalRepresentation LearningRetrievalScene Recognition+1

Fusion of Domain-Adapted Vision and Language Models for Medical Visual Question Answering

2024-04-24 · Cuong Nhat Ha, Shima Asaadi, Sanjeev Kumar Karn, Oladimeji Farri 외

Vision-language models, while effective in general domains and showing strong performance in diverse multi-modal applications like visual question-answering (VQA), struggle to maintain the same level of effectiveness in …

Language ModelingLanguage ModellingMedical Visual Question AnsweringQuestion Answering+2

OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs

2026-03-12 · Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi, Delyan Boychev 외 arxiv

Vision-Language Models (VLMs) adapted to remote sensing rely heavily on domain-specific image-text supervision, yet high-quality annotations for satellite and aerial imagery remain scarce and expensive to produce. Prevai…

Domain Adaptation