paper-with-me

Papers

AgriGPT-VL: Agricultural Vision-Language Understanding Suite

2025-10-05 · Bo Yang, Yunkui Chen, Lanfei Feng, Yu Zhang, Xiao Xu, Jianyu Zhang, Nueraili Aierken, Runhe Huang, Hongjian Lin, Yibin Ying, Shijian Li arxiv

Despite rapid advances in multimodal large language models, agricultural applications remain constrained by the scarcity of domain-tailored models, curated vision-language corpora, and rigorous evaluation. To address these challenges, we present the AgriGPT-VL Suite, a unified multimodal framework for agriculture. Our contributions are threefold. First, we introduce Agri-3M-VL, the largest vision-language corpus for agriculture to our knowledge, curated by a scalable multi-agent data generator; it comprises 1M image-caption pairs, 2M image-grounded VQA pairs, 50K expert-level VQA instances, and 15K GRPO reinforcement learning samples. Second, we develop AgriGPT-VL, an agriculture-specialized vision-language model trained via a progressive curriculum of textual grounding, multimodal shallow/deep alignment, and GRPO refinement. This method achieves strong multimodal reasoning while preserving text-only capability. Third, we establish AgriBench-VL-4K, a compact yet challenging evaluation suite with open-ended and image-grounded questions, paired with multi-metric evaluation and an LLM-as-a-judge framework. Experiments show that AgriGPT-VL outperforms leading general-purpose VLMs on AgriBench-VL-4K, achieving higher pairwise win rates in the LLM-as-a-judge evaluation. Meanwhile, it remains competitive on the text-only AgriBench-13K with no noticeable degradation of language ability. Ablation studies further confirm consistent gains from our alignment and GRPO refinement stages. We will open source all of the resources to support reproducible research and deployment in low-resource agricultural settings.

📄 PDF Abstract BibTeX arXiv:2510.04002

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMultimodal Reasoning

Similar Papers 제목 키워드 기반

AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence

2025-12-11 · Bo Yang, Lanfei Feng, Yunkui Chen, Yu Zhang 외 arxiv

Despite rapid advances in multimodal large language models, agricultural applications remain constrained by the lack of multilingual speech data, unified multimodal architectures, and comprehensive evaluation benchmarks.…

Reinforcement LearningMultimodal Reasoning

AgriGPT: a Large Language Model Ecosystem for Agriculture

2025-08-12 · Bo Yang, Yu Zhang, Lanfei Feng, Yunkui Chen 외 arxiv

Despite the rapid progress of Large Language Models (LLMs), their application in agriculture remains limited due to the lack of domain-specific models, curated datasets, and robust evaluation frameworks. To address these…

Domain Adaptation

SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation

2025-10-16 · Xiaobei Zhao, Xingqi Lyu, Xiang Li arxiv

Agricultural robots are emerging as powerful assistants across a wide range of agricultural tasks, nevertheless, they are still heavily relying on manual operations or fixed railways for movement. The A2A benchmark and t…

3D Reconstruction

Harnessing Large Vision and Language Models in Agriculture: A Review

2024-07-29 · Hongyan Zhu, Shuai Qin, Min Su, Chengzhi Lin 외

Large models can play important roles in many domains. Agriculture is another key factor affecting the lives of people around the world. It provides food, fabric, and coal for humanity. However, facing many challenges su…

Language ModellingLarge Language ModelQuestion Answering

AgroBench: Vision-Language Model Benchmark in Agriculture

2025-07-28 · Risa Shinoda, Nakamasa Inoue, Hirokatsu Kataoka, Masaki Onishi 외 arxiv

Precise automated understanding of agricultural tasks such as disease identification is essential for sustainable crop production. Recent advances in vision-language models (VLMs) are expected to further expand the range…