paper-with-me

홈 › Papers

AgriChat: A Multimodal Large Language Model for Agriculture Image Understanding

2026-03-14 · Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed arxiv

The deployment of Multimodal Large Language Models (MLLMs) in agriculture is currently stalled by a critical trade-off: the existing literature lacks the large-scale agricultural datasets required for robust model development and evaluation, while current state-of-the-art models lack the verified domain expertise necessary to reason across diverse taxonomies. To address these challenges, we propose the Vision-to-Verified-Knowledge (V2VK) pipeline, a novel generative AI-driven annotation framework that integrates visual captioning with web-augmented scientific retrieval to autonomously generate the AgriMM benchmark, effectively eliminating biological hallucinations by grounding training data in verified phytopathological literature. The AgriMM benchmark contains over 3,000 agricultural classes and more than 607k VQAs spanning multiple tasks, including fine-grained plant species identification, plant disease symptom recognition, crop counting, and ripeness assessment. Leveraging this verifiable data, we present AgriChat, a specialized MLLM that presents broad knowledge across thousands of agricultural classes and provides detailed agricultural assessments with extensive explanations. Extensive evaluation across diverse tasks, datasets, and evaluation conditions reveals both the capabilities and limitations of current agricultural MLLMs, while demonstrating AgriChat's superior performance over other open-source models, including internal and external benchmarks. The results validate that preserving visual detail combined with web-verified knowledge constitutes a reliable pathway toward robust and trustworthy agricultural AI. The code and dataset are publicly available at https://github.com/boudiafA/AgriChat .

📄 PDF Abstract BibTeX arXiv:2603.16934

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models

2024-11-30 · Yutong Zhou, Masahiro Ryo

We introduce AgriBench, the first agriculture benchmark designed to evaluate MultiModal Large Language Models (MM-LLMs) for agriculture applications. To further address the agriculture knowledge-based dataset limitation …

AgroGPT: Efficient Agricultural Vision-Language Model with Expert Tuning

2024-10-10 · Muhammad Awais, Ali Husain Salem Abdulla Alharthi, Amandeep Kumar, Hisham Cholakkal 외

Significant progress has been made in advancing large multimodal conversational models (LMMs), capitalizing on vast repositories of image-text data available online. Despite this progress, these models often encounter su…

Language ModelingLanguage Modelling

AgriGPT-VL: Agricultural Vision-Language Understanding Suite

2025-10-05 · Bo Yang, Yunkui Chen, Lanfei Feng, Yu Zhang 외 arxiv

Despite rapid advances in multimodal large language models, agricultural applications remain constrained by the scarcity of domain-tailored models, curated vision-language corpora, and rigorous evaluation. To address the…

Reinforcement LearningMultimodal Reasoning

Harnessing Large Vision and Language Models in Agriculture: A Review

2024-07-29 · Hongyan Zhu, Shuai Qin, Min Su, Chengzhi Lin 외

Large models can play important roles in many domains. Agriculture is another key factor affecting the lives of people around the world. It provides food, fabric, and coal for humanity. However, facing many challenges su…

Language ModellingLarge Language ModelQuestion Answering

Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding

2025-02-14 · Thanh-Dat Truong, Hoang-Quan Nguyen, Xuan-Bac Nguyen, Ashley Dowling 외

Multimodal conversational generative AI has shown impressive capabilities in various vision and language understanding through learning massive text-image data. However, current conversational models still lack knowledge…

General KnowledgeQuestion AnsweringSelf-Supervised Learning