paper-with-me

홈 › Papers

Seeing in Words: Learning to Classify through Language Bottlenecks

2023-06-29 · Khalid Saifullah, Yuxin Wen, Jonas Geiping, Micah Goldblum, Tom Goldstein

Neural networks for computer vision extract uninterpretable features despite achieving high accuracy on benchmarks. In contrast, humans can explain their predictions using succinct and intuitive descriptions. To incorporate explainability into neural networks, we train a vision model whose feature representations are text. We show that such a model can effectively classify ImageNet images, and we discuss the challenges we encountered when training it.

📄 PDF Abstract BibTeX arXiv:2307.00028

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Visually grounded cross-lingual keyword spotting in speech

2018-06-13 · Herman Kamper, Michael Roth

Recent work considered how images paired with speech can be used as supervision for building speech systems when transcriptions are not available. We ask whether visual grounding can be used for cross-lingual keyword spo…

Keyword SpottingVisual Grounding

It's Only Words And Words Are All I Have

2019-01-16 · Barman Manash Pratim, Dahekar Kavish, Anshuman Abhinav, Awekar Amit

The central idea of this paper is to demonstrate the strength of lyrics for music mining and natural language processing (NLP) tasks using the distributed representation paradigm. For music mining, we address two predict…

AllBinary ClassificationGeneral ClassificationMulti-class Classification

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm

2026-04-22 · Karan Goyal arxiv

The rapid proliferation of Vision-Language Models (VLMs) is often framed as enabling unified multimodal knowledge discovery but rests on an under-examined assumption: that current VLMs faithfully synthesise multimodal da…

Multimodal Reasoning

Seeing without Pixels: Perception from Camera Trajectories

2025-11-26 · Zihui Xue, Kristen Grauman, Dima Damen, Andrew Zisserman 외 arxiv

Can one perceive a video's content without seeing its pixels, just from the camera trajectory-the path it carves through space? This paper is the first to systematically investigate this seemingly implausible question. T…

Camera Pose EstimationContrastive Learning

Seeing Symbols, Missing Cultures: Probing Vision-Language Models' Reasoning on Fire Imagery and Cultural Meaning

2025-09-27 · Haorui Yu, Yang Zhao, Yijia Chu, Qiufeng Yi arxiv

Vision-Language Models (VLMs) often appear culturally competent but rely on superficial pattern matching rather than genuine cultural understanding. We introduce a diagnostic framework to probe VLM reasoning on fire-them…