paper-with-me

홈 › Papers

IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection

2025-10-16 · Zewen Li, Zitong Yu, Qilang Ye, Weicheng Xie, Wei Zhuo, Linlin Shen arxiv

The robust causal capability of Multimodal Large Language Models (MLLMs) hold the potential of detecting defective objects in Industrial Anomaly Detection (IAD). However, most traditional IAD methods lack the ability to provide multi-turn human-machine dialogues and detailed descriptions, such as the color of objects, the shape of an anomaly, or specific types of anomalies. At the same time, methods based on large pre-trained models have not fully stimulated the ability of large models in anomaly detection tasks. In this paper, we explore the combination of rich text semantics with both image-level and pixel-level information from images and propose IAD-GPT, a novel paradigm based on MLLMs for IAD. We employ Abnormal Prompt Generator (APG) to generate detailed anomaly prompts for specific objects. These specific prompts from the large language model (LLM) are used to activate the detection and segmentation functions of the pre-trained visual-language model (i.e., CLIP). To enhance the visual grounding ability of MLLMs, we propose Text-Guided Enhancer, wherein image features interact with normal and abnormal text prompts to dynamically select enhancement pathways, which enables language models to focus on specific aspects of visual data, enhancing their ability to accurately interpret and respond to anomalies within images. Moreover, we design a Multi-Mask Fusion module to incorporate mask as expert knowledge, which enhances the LLM's perception of pixel-level anomalies. Extensive experiments on MVTec-AD and VisA datasets demonstrate our state-of-the-art performance on self-supervised and few-shot anomaly detection and segmentation tasks, such as MVTec-AD and VisA datasets. The codes are available at \href{https://github.com/LiZeWen1225/IAD-GPT}{https://github.com/LiZeWen1225/IAD-GPT}.

📄 PDF Abstract BibTeX arXiv:2510.16036

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly DetectionVisual Grounding

Similar Papers 제목 키워드 기반

Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment

2024-02-21 · Yunxin Li, Xinyu Chen, Baotian Hu, Haoyuan Shi 외

Evaluating and Rethinking the current landscape of Large Multimodal Models (LMMs), we observe that widely-used visual-language projection approaches (e.g., Q-former or MLP) focus on the alignment of image-text descriptio…

Language ModellingQuestion AnsweringSmall Language ModelVisual Question Answering+1

EchoSight: Advancing Visual-Language Models with Wiki Knowledge

2024-07-17 · Yibin Yan, Weidi Xie

Knowledge-based Visual Question Answering (KVQA) tasks require answering questions about images using extensive background knowledge. Despite significant advancements, generative models often struggle with these tasks du…

ArticlesQuestion AnsweringRAGRetrieval+3

VisualQuest: A Diverse Image Dataset for Evaluating Visual Recognition in LLMs

2025-03-25 · Kelaiti Xiao, Liang Yang, Paerhati Tulajiang, Hongfei Lin

This paper introduces VisualQuest, a novel image dataset designed to assess the ability of large language models (LLMs) to interpret non-traditional, stylized imagery. Unlike conventional photographic benchmarks, VisualQ…

DiversityMultimodal Reasoning

M$^3$-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering

2026-04-28 · Jiatong Ma, Longteng Guo, Yuchen Liu, Zijia Zhao 외 arxiv

We present M$^3$-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine-grained multimodal entity understanding and complex m…

Visual Question AnsweringMultimodal Reasoning

OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval

2025-05-10 · Wei Yang, Jingjing Fu, Rui Wang, Jinyu Wang 외

Vision-language retrieval-augmented generation (RAG) has become an effective approach for tackling Knowledge-Based Visual Question Answering (KB-VQA), which requires external knowledge beyond the visual content presented…

Cross-Modal RetrievalQuestion AnsweringRAGReranking+4