paper-with-me

홈 › Papers

AgriDoctor: A Multimodal Intelligent Assistant for Agriculture

2025-09-21 · Mingqing Zhang, Zhuoning Xu, Peijie Wang, Rongji Li, Liang Wang, Qiang Liu, Jian Xu, Xuyao Zhang, Shu Wu, Liang Wang arxiv

Accurate crop disease diagnosis is essential for sustainable agriculture and global food security. Existing methods, which primarily rely on unimodal models such as image-based classifiers and object detectors, are limited in their ability to incorporate domain-specific agricultural knowledge and lack support for interactive, language-based understanding. Recent advances in large language models (LLMs) and large vision-language models (LVLMs) have opened new avenues for multimodal reasoning. However, their performance in agricultural contexts remains limited due to the absence of specialized datasets and insufficient domain adaptation. In this work, we propose AgriDoctor, a modular and extensible multimodal framework designed for intelligent crop disease diagnosis and agricultural knowledge interaction. As a pioneering effort to introduce agent-based multimodal reasoning into the agricultural domain, AgriDoctor offers a novel paradigm for building interactive and domain-adaptive crop health solutions. It integrates five core components: a router, classifier, detector, knowledge retriever and LLMs. To facilitate effective training and evaluation, we construct AgriMM, a comprehensive benchmark comprising 400000 annotated disease images, 831 expert-curated knowledge entries, and 300000 bilingual prompts for intent-driven tool selection. Extensive experiments demonstrate that AgriDoctor, trained on AgriMM, significantly outperforms state-of-the-art LVLMs on fine-grained agricultural tasks, establishing a new paradigm for intelligent and sustainable farming applications.

📄 PDF Abstract BibTeX arXiv:2509.17044

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningDomain Adaptation

Similar Papers 제목 키워드 기반

GigaPevt: Multimodal Medical Assistant

2024-02-26 · Pavel Blinov, Konstantin Egorov, Ivan Sviridov, Nikolay Ivanov 외

Building an intelligent and efficient medical assistant is still a challenging AI problem. The major limitation comes from the data modality scarceness, which reduces comprehensive patient perception. This demo paper pre…

Question Answering

Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis

2024-06-17 · Kubilay Can Demir, Belen Lojo Rodriguez, Tobias Weise, Andreas Maier 외

To develop intelligent speech assistants and integrate them seamlessly with intra-operative decision-support frameworks, accurate and efficient surgical phase recognition is a prerequisite. In this study, we propose a mu…

Surgical phase recognition

Chatbot Application to Support Smart Agriculture in Thailand

2023-07-31 · Paweena Suebsombut, Pradorn Sureephong, Aicha Sekhari, Suepphong Chernbumroong 외

A chatbot is a software developed to help reply to text or voice conversations automatically and quickly in real time. In the agriculture sector, the existing smart agriculture systems just use data from sensing and inte…

ChatbotDecision MakingRecommendation Systems

Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases

2024-12-03 · Liqiong Wang, Teng Jin, Jinyu Yang, Ales Leonardis 외

In the general domain, large multimodal models (LMMs) have achieved significant advancements, yet challenges persist in applying them to specific fields, especially agriculture. As the backbone of the global economy, agr…

Instruction Following

HUMBO: Bridging Response Generation and Facial Expression Synthesis

2019-05-24 · Shang-Yu Su, Po-Wei Lin, Yun-Nung Chen

Spoken dialogue systems that assist users to solve complex tasks such as movie ticket booking have become an emerging research topic in artificial intelligence and natural language processing areas. With a well-designed …

Dialogue Generationmultimodal interactionResponse GenerationSpoken Dialogue Systems