paper-with-me

홈 › Papers

SAViR-T: Spatially Attentive Visual Reasoning with Transformers

2022-06-18 · Pritish Sahu, Kalliopi Basioti, Vladimir Pavlovic

We present a novel computational model, "SAViR-T", for the family of visual reasoning problems embodied in the Raven's Progressive Matrices (RPM). Our model considers explicit spatial semantics of visual elements within each image in the puzzle, encoded as spatio-visual tokens, and learns the intra-image as well as the inter-image token dependencies, highly relevant for the visual reasoning task. Token-wise relationship, modeled through a transformer-based SAViR-T architecture, extract group (row or column) driven representations by leveraging the group-rule coherence and use this as the inductive bias to extract the underlying rule representations in the top two row (or column) per token in the RPM. We use this relation representations to locate the correct choice image that completes the last row or column for the RPM. Extensive experiments across both synthetic RPM benchmarks, including RAVEN, I-RAVEN, RAVEN-FAIR, and PGM, and the natural image-based "V-PROM" demonstrate that SAViR-T sets a new state-of-the-art for visual reasoning, exceeding prior models' performance by a considerable margin.

📄 PDF Abstract BibTeX arXiv:2206.09265

Code (1)

kalbasioti/visual-reasoning 공식 구현 pytorch

Tasks

Inductive BiasVisual Reasoning

Methods 이 논문이 사용한 방법론

PGM A regularization criterion that, differently from dropout and its variants, is deterministic rather than random. It grounds on the…

Similar Papers 제목 키워드 기반

Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers

2026-03-03 · Youngjun Jun, Seil Kang, Woojung Han, Seong Jae Hwang arxiv

Video Diffusion Transformers (DiTs) have been synthesizing high-quality video with high fidelity from given text descriptions involving motion. However, understanding how Video DiTs convert motion words into video remain…

Video Semantic Segmentation

Exploring Object-Aware Attention Guided Frame Association for RGB-D SLAM

2022-01-28 · Ali Caglayan, Nevrez Imamoglu, Oguzhan Guclu, Ali Osman Serhatoglu 외

Deep learning models as an emerging topic have shown great progress in various fields. Especially, visualization tools such as class activation mapping methods provided visual explanation on the reasoning of convolutiona…

ObjectSimultaneous Localization and Mapping

Exploring Object-Aware Attention Guided Frame Association for RGB-D SLAM

2025-10-30 · Ali Caglayan, Nevrez Imamoglu, Oguzhan Guclu, Ali Osman Serhatoglu 외 arxiv

Attention models have recently emerged as a powerful approach, demonstrating significant progress in various fields. Visualization techniques, such as class activation mapping, provide visual insights into the reasoning …

Spatially Aware Multimodal Transformers for TextVQA

2020-07-23 · ECCV 2020 8 · Yash Kant, Dhruv Batra, Peter Anderson, Alex Schwing 외

Textual cues are essential for everyday tasks like buying groceries and using public transport. To develop this assistive technology, we study the TextVQA task, i.e., reasoning about text in images to answer a question. …

Optical Character Recognition (OCR)Spatial ReasoningTextVQAVisual Grounding+1

BiLSTM-VHP: BiLSTM-Powered Network for Viral Host Prediction

2025-09-14 · Azher Ahmed Efat, Farzana Islam, Annajiat Alim Rasel, Munima Haque arxiv

Recorded history shows the long coexistence of humans and animals, suggesting it began much earlier. Despite some beneficial interdependence, many animals carry viral diseases that can spread to humans. These diseases ar…