paper-with-me

Papers

Spatial Knowledge Distillation to aid Visual Reasoning

2018-12-10 · Somak Aditya, Rudra Saha, Yezhou Yang, Chitta Baral

For tasks involving language and vision, the current state-of-the-art methods tend not to leverage any additional information that might be present to gather relevant (commonsense) knowledge. A representative task is Visual Question Answering where large diagnostic datasets have been proposed to test a system's capability of answering questions about images. The training data is often accompanied by annotations of individual object properties and spatial locations. In this work, we take a step towards integrating this additional privileged information in the form of spatial knowledge to aid in visual reasoning. We propose a framework that combines recent advances in knowledge distillation (teacher-student framework), relational reasoning and probabilistic logical languages to incorporate such knowledge in existing neural networks for the task of Visual Question Answering. Specifically, for a question posed against an image, we use a probabilistic logical language to encode the spatial knowledge and the spatial understanding about the question in the form of a mask that is directly provided to the teacher network. The student network learns from the ground-truth information as well as the teachers prediction via distillation. We also demonstrate the impact of predicting such a mask inside the teachers network using attention. Empirically, we show that both the methods improve the test accuracy over a state-of-the-art approach on a publicly available dataset.

📄 PDF Abstract BibTeX arXiv:1812.03631

Code (0)

등록된 구현이 없습니다.

Tasks

DiagnosticKnowledge DistillationQuestion AnsweringRelational ReasoningVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

Similar Papers 제목 키워드 기반

Dense 2D-3D Indoor Prediction with Sound via Aligned Cross-Modal Distillation

2023-09-20 · ICCV 2023 1 · Heeseung Yun, Joonil Na, Gunhee Kim

Sound can convey significant information for spatial reasoning in our daily lives. To endow deep networks with such ability, we address the challenge of dense indoor prediction with sound in both 2D and 3D via cross-moda…

3D Scene ReconstructionDepth EstimationKnowledge DistillationPrediction+3

GLaD: Geometric Latent Distillation for Vision-Language-Action Models

2025-12-10 · Minghao Guo, Meng Cao, Jiachen Tao, Rongtao Xu 외 arxiv

Most existing Vision-Language-Action (VLA) models rely primarily on RGB information, while ignoring geometric cues crucial for spatial reasoning and manipulation. In this work, we introduce GLaD, a geometry-aware VLA fra…

Knowledge DistillationSpatial Reasoning

SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning

2025-05-18 · Yang Liu, Ming Ma, Xiaomin Yu, Pengxiang Ding 외

Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for integrating spatial cues, such as point clou…

Knowledge DistillationSpatial Reasoning

Geospatial-Reasoning-Driven Vocabulary-Agnostic Remote Sensing Semantic Segmentation

2026-02-09 · Chufeng Zhou, Jian Wang, Xinyuan Liu, Xiaokang Zhang arxiv

Open-vocabulary semantic segmentation has become an important direction in remote sensing, as it enables recognition beyond predefined land-cover categories. However, existing methods mainly depend on passive visual-text…

Knowledge DistillationSemantic Segmentation

AVQACL: A Novel Benchmark for Audio-Visual Question Answering Continual Learning

2025-01-01 · CVPR 2025 1 · Kaixuan Wu, Xinde Li, Xinling Li, Chuanfei Hu 외

In this paper, a novel benchmark for audio-visual question answering continual learning (AVQACL) is introduced, aiming to study fine-grained scene understanding and spatial-temporal reasoning in videos under a contin…

Audio-visual Question AnsweringContinual LearningKnowledge DistillationQuestion Answering+2