paper-with-me

홈 › Papers

Spatially Grounded Concept-Based Image Classification

2025-10-05 · Ran Eisenberg, Amit Rozner, Ethan Fetaya, Ofir Lindenbaum arxiv

Deep neural networks can achieve high accuracy while relying on evidence that is hard to inspect or misaligned with the intended task. Concept Bottleneck Models (CBMs) expose human-interpretable concepts, but most treat concepts as global attributes and do not show how localized evidence is aggregated into a decision. We propose \textbf{SEG-MIL-CBM}, a spatially grounded CBM that decomposes each image into concept-guided regions and classifies it by attention-based aggregation of segment-level concept evidence. The same segment evidence terms form the prediction and the explanation, exposing which regions and concepts support the predicted logit without a separate post-hoc attribution module. Among evaluated CBM-family baselines, SEG-MIL-CBM improves Waterbirds worst-group accuracy from $65.1\%$ to $72.0\%$, reaches $87.4\%$ worst-group accuracy on Pawrious, remains competitive on standard recognition, and attains the best CBM accuracy on CIFAR-100 ($85.3\%$). Segment-level faithfulness experiments on CUB further show that its learned segment ranking matches or improves over evaluated segment-ranking controls.

📄 PDF Abstract BibTeX arXiv:2510.04180

Code (0)

등록된 구현이 없습니다.

Tasks

Image Classification

Results from the Paper

RankTaskDatasetModelMetrics
#2 Classification CIFAR-100 Spatially Grounded Concept-Based Image C Accuracy: 85.3
#65 Image Classification CIFAR-100 Spatially Grounded Concept-Based Image C Percentage correct: 85.3

Similar Papers 제목 키워드 기반

CFM: Language-aligned Concept Foundation Model for Vision

2026-01-20 · Kai Wittenmayer, Sukrut Rao, Amin Parchami-Araghi, Bernt Schiele 외 arxiv

Language-aligned vision foundation models perform strongly across diverse downstream tasks. Yet, their learned representations remain opaque, making interpreting their decision-making difficult. Recent work decompose the…

Image Classification

Show and Tell: Visually Explainable Deep Neural Nets via Spatially-Aware Concept Bottleneck Models

2025-02-27 · CVPR 2025 1 · Itay Benou, Tammy Riklin-Raviv

Modern deep neural networks have now reached human-level performance across a variety of tasks. However, unlike humans they lack the ability to explain their decisions by showing where and telling what concepts guided th…

Zero Shot Segmentation

Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI

2026-04-17 · Lama Moukheiber, Caleb M. Yeung, Haotian Xue, Alec Helbling 외 arxiv

Spatial reasoning and visual grounding are core capabilities for vision-language models (VLMs), yet most medical VLMs produce predictions without transparent reasoning or spatial evidence. Existing benchmarks also evalua…

Visual Question AnsweringSpatial ReasoningVisual Grounding

Spatially Grounded Long-Horizon Task Planning in the Wild

2026-03-13 · Sehun Jung, HyunJee Song, Dong-Hee Kim, Reuben Tan 외 arxiv

Recent advances in robot manipulation increasingly leverage Vision-Language Models (VLMs) for high-level reasoning, such as decomposing task instructions into sequential action plans expressed in natural language that gu…

Robot Manipulation

GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks

2025-12-31 · Wei Zhang, Chaoqun Wang, Zixuan Guan, Sam Kao 외 arxiv

Systematic failures of vision models on semantically coherent subsets, known as error slices, reveal limitations in robustness and evaluation. Existing slice discovery approaches largely model slices as clusters in repre…

Object Detection