paper-with-me

홈 › Papers

Empowering CAM-Based Methods with Capability to Generate Fine-Grained and High-Faithfulness Explanations

2023-03-16 · Changqing Qiu, Fusheng Jin, Yining Zhang

Recently, the explanation of neural network models has garnered considerable research attention. In computer vision, CAM (Class Activation Map)-based methods and LRP (Layer-wise Relevance Propagation) method are two common explanation methods. However, since most CAM-based methods can only generate global weights, they can only generate coarse-grained explanations at a deep layer. LRP and its variants, on the other hand, can generate fine-grained explanations. But the faithfulness of the explanations is too low. To address these challenges, in this paper, we propose FG-CAM (Fine-Grained CAM), which extends CAM-based methods to enable generating fine-grained and high-faithfulness explanations. FG-CAM uses the relationship between two adjacent layers of feature maps with resolution differences to gradually increase the explanation resolution, while finding the contributing pixels and filtering out the pixels that do not contribute. Our method not only solves the shortcoming of CAM-based methods without changing their characteristics, but also generates fine-grained explanations that have higher faithfulness than LRP and its variants. We also present FG-CAM with denoising, which is a variant of FG-CAM and is able to generate less noisy explanations with almost no change in explanation faithfulness. Experimental results show that the performance of FG-CAM is almost unaffected by the explanation resolution. FG-CAM outperforms existing CAM-based methods significantly in both shallow and intermediate layers, and outperforms LRP and its variants significantly in the input layer. Our code is available at https://github.com/dongmo-qcq/FG-CAM.

📄 PDF Abstract BibTeX arXiv:2303.09171

Code (1)

dongmo-qcq/fg-cam 공식 구현 pytorch

Tasks

Explanation Generation

Methods 이 논문이 사용한 방법론

CAM Class activation maps could be used to interpret the prediction decision made by the convolutional neural network (CNN). Image source: [Learning Deep Features for…

Similar Papers 제목 키워드 기반

FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models

2024-07-16 · Pengxiang Li, Zhi Gao, Bofei Zhang, Tao Yuan 외

Vision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn…

RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation

2024-12-16 · Xiaoxi Li, Jiajie Jin, Yujia Zhou, Yongkang Wu 외

Large language models (LLMs) exhibit remarkable generative capabilities but often suffer from hallucinations. Retrieval-augmented generation (RAG) offers an effective solution by incorporating external knowledge, but exi…

RAGRetrievalRetrieval-augmented Generation

Small Agent Can Also Rock! Empowering Small Language Models as Hallucination Detector

2024-06-17 · Xiaoxue Cheng, Junyi Li, Wayne Xin Zhao, Hongzhi Zhang 외

Hallucination detection is a challenging task for large language models (LLMs), and existing studies heavily rely on powerful closed-source LLMs such as GPT-4. In this paper, we propose an autonomous LLM-based agent fram…

2kHallucination

Follow the Attention: Combining Partial Pose and Object Motion for Fine-Grained Action Detection

2019-05-11 · Mohammad Mahdi Kazemi Moghaddam, Ehsan Abbasnejad, Javen Shi

Retailers have long been searching for ways to effectively understand their customers' behaviour in order to provide a smooth and pleasant shopping experience that attracts more customers everyday and maximises their rev…

Action DetectionActivity DetectionActivity RecognitionFine-Grained Action Detection+1

Empowering Reliable Visual-Centric Instruction Following in MLLMs

2026-01-06 · Weilei He, Feng Ju, Zhiyuan Fan, Rui Min 외 arxiv

Evaluating the instruction-following (IF) capabilities of Multimodal Large Language Models (MLLMs) is essential for rigorously assessing how faithfully model outputs adhere to user-specified intentions. Nevertheless, exi…

Instruction Following