paper-with-me

홈 › Papers

Vision-Language Semantic Grounding for Multi-Domain Crop-Weed Segmentation

2026-02-27 · Nazia Hossain, Xintong Jiang, Yu Tian, Philippe Seguin, O. Grant Clark, Shangpeng Sun arxiv

Fine-grained crop-weed segmentation is essential for enabling targeted herbicide application in precision agriculture. However, existing deep learning models struggle to generalize across heterogeneous agricultural environments due to reliance on dataset-specific visual features. We propose Vision-Language Weed Segmentation (VL-WS), a novel framework that addresses this limitation by grounding pixel-level segmentation in semantically aligned, domain-invariant representations. Our architecture employs a dual-encoder design, where frozen Contrastive Language-Image Pretraining (CLIP) embeddings and task-specific spatial features are fused and modulated via Feature-wise Linear Modulation (FiLM) layers conditioned on natural language captions. This design enables image level textual descriptions to guide channel-wise feature refinement while preserving fine-grained spatial localization. Unlike prior works restricted to training and evaluation on single-source datasets, VL-WS is trained on a unified corpus that includes close-range ground imagery (robotic platforms) and high-altitude UAV imagery, covering diverse crop types, weed species, growth stages, and sensing conditions. Experimental results across four benchmark datasets demonstrate the effectiveness of our framework, with VL-WS achieving a mean Dice score of 91.64% and outperforming the CNN baseline by 4.98%. The largest gains occur on the most challenging weed class, where VL-WS attains 80.45% Dice score compared to 65.03% for the best baseline, representing a 15.42% improvement. VL-WS further maintains stable weed segmentation performance under limited target-domain supervision, indicating improved generalization and data efficiency. These findings highlight the potential of vision-language alignment to enable scalable, label-efficient segmentation models deployable across diverse real-world agricultural domains.

📄 PDF Abstract BibTeX arXiv:2602.23677

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Can Feedback Enhance Semantic Grounding in Large Vision-Language Models?

2024-04-09 · Yuan-Hong Liao, Rafid Mahmood, Sanja Fidler, David Acuna

Enhancing semantic grounding abilities in Vision-Language Models (VLMs) often involves collecting domain-specific training data, refining the network architectures, or modifying the training recipes. In this work, we ven…

Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models

2023-09-07 · Jiaying Lu, Jinmeng Rao, Kezhen Chen, Xiaoyuan Guo 외

Large Vision-Language Models (LVLMs) offer remarkable benefits for a variety of vision-language tasks. However, a challenge hindering their application in real-world scenarios, particularly regarding safety, robustness, …

Question AnsweringVisual Question Answering

Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?

2025-01-01 · CVPR 2025 1 · Yuan-Hong Liao, Rafid Mahmood, Sanja Fidler, David Acuna

Improving semantic grounding in Vision-Language Models (VLMs) often involves collecting domain-specific training data, refining the network architectures, or modifying the training recipes. In this work, we venture i…

HierVL: Semi-Supervised Segmentation leveraging Hierarchical Vision-Language Synergy with Dynamic Text-Spatial Query Alignment

2025-06-16 · Numair Nadeem, Saeed Anwar, Muhammad Hamza Asad, Abdul Bais

Semi-supervised semantic segmentation remains challenging under severe label scarcity and domain variability. Vision-only methods often struggle to generalize, resulting in pixel misclassification between similar classes…

Semantic SegmentationSemi-Supervised Semantic Segmentation

Meanings and Measurements: Multi-Agent Probabilistic Grounding for Vision-Language Navigation

2026-03-19 · Swagat Padhan, Lakshya Jain, Bhavya Minesh Shah, Omkar Patil 외 arxiv

Robots collaborating with humans must convert natural language goals into actionable, physically grounded decisions. For example, executing a command such as "go two meters to the right of the fridge" requires grounding …

Vision-Language Navigation