paper-with-me

홈 › Papers

Reading Is Believing: Revisiting Language Bottleneck Models for Image Classification

2024-06-22 · Honori Udo, Takafumi Koshinaka

We revisit language bottleneck models as an approach to ensuring the explainability of deep learning models for image classification. Because of inevitable information loss incurred in the step of converting images into language, the accuracy of language bottleneck models is considered to be inferior to that of standard black-box models. Recent image captioners based on large-scale foundation models of Vision and Language, however, have the ability to accurately describe images in verbal detail to a degree that was previously believed to not be realistically possible. In a task of disaster image classification, we experimentally show that a language bottleneck model that combines a modern image captioner with a pre-trained language model can achieve image classification accuracy that exceeds that of black-box models. We also demonstrate that a language bottleneck model and a black-box model may be thought to extract different features from images and that fusing the two can create a synergistic effect, resulting in even higher classification accuracy.

📄 PDF Abstract BibTeX arXiv:2406.15816

Code (0)

등록된 구현이 없습니다.

Tasks

Classificationimage-classificationImage ClassificationLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Reading Isn't Believing: Adversarial Attacks On Multi-Modal Neurons

2021-03-18 · David A. Noever, Samantha E. Miller Noever

With Open AI's publishing of their CLIP model (Contrastive Language-Image Pre-training), multi-modal neural networks now provide accessible models that combine reading with visual recognition. Their network offers novel …

Base-based Model Checking for Multi-Agent Only Believing (long version)

2023-07-27 · Tiago de Lima, Emiliano Lorini, François Schwarzentruber

We present a novel semantics for the language of multi-agent only believing exploiting belief bases, and show how to use it for automatically checking formulas of this language and of its dynamic extension with private b…

MoVie: Revisiting Modulated Convolutions for Visual Counting and Beyond

2020-04-24 · ICLR 2021 1 · Duy-Kien Nguyen, Vedanuj Goswami, Xinlei Chen

This paper focuses on visual counting, which aims to predict the number of occurrences given a natural image and a query (e.g. a question or a category). Unlike most prior works that use explicit, symbolic models which c…

Object CountingQuestion AnsweringVisual Question Answering (VQA)

Robust Place Recognition using an Imaging Lidar

2021-03-03 · Tixiao Shan, Brendan Englot, Fabio Duarte, Carlo Ratti 외

We propose a methodology for robust, real-time place recognition using an imaging lidar, which yields image-quality high-resolution 3D point clouds. Utilizing the intensity readings of an imaging lidar, we project the po…

Transformer-based Language Model Fine-tuning Methods for COVID-19 Fake News Detection

2021-01-14 · Ben Chen, Bin Chen, Dehong Gao, Qijin Chen 외

With the pandemic of COVID-19, relevant fake news is spreading all over the sky throughout the social media. Believing in them without discrimination can cause great trouble to people's life. However, universal language …

Fake News DetectionLanguage ModelingLanguage Modelling