paper-with-me

홈 › Papers

Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks

2026-05-26 · Partho Ghose, Al Bashir, Prem Raj, Azlan Zahid arxiv

Large Language Models (LLMs) are being rapidly adopted in agricultural imaging applications, ranging from crop interpretation to synthetic field image generation. However, these models frequently exhibit hallucinations outputs that appear confident yet deviate from biological or environmental reality potentially leading to misinformed agronomic insights. This study investigates such hallucinations in two complementary directions: image-to-text, where LLMs interpret crop or field imagery to describe conditions such as biotic and abiotic stresses, and text-to-image, where models generate synthetic agricultural scenes based on descriptive prompts. We examine errors involving biological inconsistency, contextual inaccuracy, and agronomic implausibility, evaluating the outputs under domain-informed criteria across multiple imaging modalities. Our analysis identifies recurring hallucination patterns within both interpretive and generative tasks. In image interpretation, LLMs (e.g., Gemma, LLAVA, Qwen, and MiniCPM) achieved modest zero-shot accuracy (63 to 75 percent), whereas few-shot prompting improved performance up to 86.8 percent, exhibiting false detections and missed infections, indicating residual hallucination effects. In text-to-image tasks, advanced models such as GPT-5 and Gemini 2.5 Flash generate up to 91 percent biologically inconsistent scenes under relaxed prompt constraints, revealing fundamental weaknesses in current LLMs. This systematic assessment of visual reasoning and generation offers critical insights toward enhancing the reliability and trustworthiness of LLM-based agricultural imaging platforms.

📄 PDF Abstract BibTeX arXiv:2605.27595

Code (0)

등록된 구현이 없습니다.

Tasks

Visual ReasoningImage Generation

Similar Papers 제목 키워드 기반

AgriChat: A Multimodal Large Language Model for Agriculture Image Understanding

2026-03-14 · Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed arxiv

The deployment of Multimodal Large Language Models (MLLMs) in agriculture is currently stalled by a critical trade-off: the existing literature lacks the large-scale agricultural datasets required for robust model develo…

RSHallu: Dual-Mode Hallucination Evaluation for Remote-Sensing Multimodal Large Language Models with Domain-Tailored Mitigation

2026-02-11 · Zihui Zhou, Yong Feng, Yanying Chen, Guofan Duan 외 arxiv

Multimodal large language models (MLLMs) are increasingly adopted in remote sensing (RS) and have shown strong performance on tasks such as RS visual grounding (RSVG), RS visual question answering (RSVQA), and multimodal…

Visual Question AnsweringVisual Grounding

AgriRegion: Region-Aware Retrieval for High-Fidelity Agricultural Advice

2025-12-10 · Mesafint Fanuel, Mahmoud Nabil Mahmoud, Crystal Cook Marshal, Vishal Lakhotia 외 arxiv

Large Language Models (LLMs) have demonstrated significant potential in democratizing access to information. However, in the domain of agriculture, general-purpose models frequently suffer from contextual hallucination, …

Semantic Similarity

Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

2024-01-19 · Xiyao Wang, YuHang Zhou, Xiaoyu Liu, Hongjin Lu 외

Multimodal Large Language Models (MLLMs) have demonstrated proficiency in handling a variety of visual-language tasks. However, current MLLM benchmarks are predominantly designed to evaluate reasoning based on static inf…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback

2023-12-01 · CVPR 2024 1 · Tianyu Yu, Yuan YAO, Haoye Zhang, Taiwen He 외

Multimodal Large Language Models (MLLMs) have recently demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction. However, existing MLLMs prevalently suffer from serious hallucination p…

HallucinationImage CaptioningVisual Question Answering