paper-with-me

홈 › Papers

Visual Similarity Attention

2019-11-18 · Meng Zheng, Srikrishna Karanam, Terrence Chen, Richard J. Radke, Ziyan Wu

While there has been substantial progress in learning suitable distance metrics, these techniques in general lack transparency and decision reasoning, i.e., explaining why the input set of images is similar or dissimilar. In this work, we solve this key problem by proposing the first method to generate generic visual similarity explanations with gradient-based attention. We demonstrate that our technique is agnostic to the specific similarity model type, e.g., we show applicability to Siamese, triplet, and quadruplet models. Furthermore, we make our proposed similarity attention a principled part of the learning process, resulting in a new paradigm for learning similarity functions. We demonstrate that our learning mechanism results in more generalizable, as well as explainable, similarity models. Finally, we demonstrate the generality of our framework by means of experiments on a variety of tasks, including image retrieval, person re-identification, and low-shot semantic segmentation.

📄 PDF Abstract BibTeX arXiv:1911.07381

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalPerson Re-IdentificationRetrievalSemantic SegmentationTriplet

Similar Papers 제목 키워드 기반

Towards Visually Explaining Similarity Models

2020-08-13 · Meng Zheng, Srikrishna Karanam, Terrence Chen, Richard J. Radke 외

We consider the problem of visually explaining similarity models, i.e., explaining why a model predicts two images to be similar in addition to producing a scalar score. While much recent work in visual model interpretab…

Image RetrievalMetric LearningPerson Re-IdentificationRetrieval+1

Interpreting Attention Models with Human Visual Attention in Machine Reading Comprehension

2020-06-03 · Anonymous

While neural networks with attention mechanisms have achieved superior performance on many natural language processing tasks, it remains unclear to which extent learned attention resembles human visual attention. We stud…

Machine Reading ComprehensionQuestion AnsweringReading Comprehension

Self-attention in Vision Transformers Performs Perceptual Grouping, Not Attention

2023-03-02 · Paria Mehrani, John K. Tsotsos

Recently, a considerable number of studies in computer vision involves deep neural architectures called vision transformers. Visual processing in these models incorporates computational models that are claimed to impleme…

Saliency Detection

Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models

2026-04-13 · Kexin Ma, Jing Xiao, Chaofeng Chen, Geyong Min 외 arxiv

Token pruning has emerged as an effective approach to reduce the substantial computational overhead of Large Vision-Language Models (LVLMs) by discarding less informative visual tokens while preserving performance. Howev…

Report: Dynamic Eye Movement Matching and Visualization Tool in Neuro Gesture

2017-12-27 · Qiangeng Xu, John Kender

In the research of the impact of gestures using by a lecturer, one challenging task is to infer the attention of a group of audiences. Two important measurements that can help infer the level of attention are eye movemen…

EEGElectroencephalogram (EEG)