paper-with-me

Papers

Unlocking Open-Set Language Accessibility in Vision Models

2025-03-14 · Fawaz Sammani, Jonas Fischer, Nikos Deligiannis

Visual classifiers offer high-dimensional feature representations that are challenging to interpret and analyze. Text, in contrast, provides a more expressive and human-friendly interpretable medium for understanding and analyzing model behavior. We propose a simple, yet powerful method for reformulating any visual classifier so that it can be accessed with open-set text queries without compromising its original performance. Our approach is label-free, efficient, and preserves the underlying classifier's distribution and reasoning processes. We thus unlock several text-based interpretability applications for any classifier. We apply our method on 40 visual classifiers and demonstrate two primary applications: 1) building both label-free and zero-shot concept bottleneck models and therefore converting any classifier to be inherently-interpretable and 2) zero-shot decoding of visual features into natural language. In both applications, we achieve state-of-the-art results, greatly outperforming existing works. Our method enables text approaches for interpreting visual classifiers.

📄 PDF Abstract BibTeX arXiv:2503.10981

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WebAccessVL: Violation-Aware VLM for Web Accessibility

2025-12-19 · Amber Yijia Zheng, Jae Joong Lee, Bedrich Benes, Raymond A. Yeh arxiv

We present a vision-language model (VLM) that automatically edits website HTML to address violations of the Web Content Accessibility Guidelines 2 (WCAG2) while preserving the original design. We formulate this as a supe…

Program Synthesis

Comprehensive Analysis of Transparency and Accessibility of ChatGPT, DeepSeek, And other SoTA Large Language Models

2025-02-21 · Ranjan Sapkota, Shaina Raza, Manoj Karkee

Despite increasing discussions on open-source Artificial Intelligence (AI), existing research lacks a discussion on the transparency and accessibility of state-of-the-art (SoTA) Large Language Models (LLMs). The Open Sou…

Domain Adaptation

MIDAL: A Dataset of Math Image Descriptions for Accessible Learning

2026-08-01 · Rebeka Popek, Vaghawan Ojha, Young Hwan You arxiv

Many open educational resources are lacking in accessibility, especially in-depth image descriptions. In subjects like Science and Mathematics, however, it can be particularly difficult to write image descriptions since …

Mathematical Reasoning

Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image

2024-07-07 · Pengkun Jiao, Na Zhao, Jingjing Chen, Yu-Gang Jiang

Open-vocabulary 3D object detection (OV-3DDet) aims to localize and recognize both seen and previously unseen object categories within any new 3D scene. While language and vision foundation models have achieved success i…

3D Object DetectionObjectobject-detectionObject Detection

Template-Based Text-to-Image Alignment for Language Accessibility: A Study on Visualizing Text Simplifications

2025-10-13 · Belkiss Souayed, Sarah Ebling, Yingqiang Gao arxiv

Individuals with intellectual disabilities often have difficulties in comprehending complex texts. While many text-to-image models prioritize aesthetics over accessibility, it is not clear how visual illustrations relate…