paper-with-me

홈 › Papers

Stacked Latent Attention for Multimodal Reasoning

2018-06-01 · CVPR 2018 6 · Haoqi Fan, Jiatong Zhou

Attention has shown to be a pivotal development in deep learning and has been used for a multitude of multimodal learning tasks such as visual question answering and image captioning. In this work, we pinpoint the potential limitations to the design of a traditional attention model. We identify that 1) current attention mechanisms discard the latent information from intermediate reasoning, losing the positional information already captured by the attention heatmaps and 2) stacked attention, a common way to improve spatial reasoning, may have suboptimal performance because of the vanishing gradient problem. We introduce a novel attention architecture to address these problems, in which all spatial configuration information contained in the intermediate reasoning process is retained in a pathway of convolutional layers. We show that this new attention leads to substantial improvements in multiple multimodal reasoning tasks, including achieving single model performance without using external knowledge comparable to the state-of-the-art on the VQA dataset, as well as clear gains for the image captioning task.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningMultimodal ReasoningQuestion AnsweringSpatial ReasoningVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models

2025-08-18 · Tan-Hanh Pham, Chris Ngo arxiv

Many reasoning techniques for large multimodal models adapt language model approaches, such as Chain-of-Thought (CoT) prompting, which express reasoning as word sequences. While effective for text, these methods are subo…

Multimodal Reasoning

Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders

2025-10-29 · Ali Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk 외 arxiv

Despite significant advances in Multimodal Large Language Models (MLLMs), understanding complex temporal dynamics in videos remains a major challenge. Our experiments show that current Video Large Language Model (Video-L…

Video Question AnsweringAction Recognition

OPLD: On-Policy Latent Distillation for Multimodal Reasoning

2026-07-30 · Shoutai Zhu, Tianyang Xu, Bin Sun, Mingyuan Xu 외 arxiv

Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However, existing approaches remain constrained by externally defined reason…

Multimodal ReasoningVisual Reasoning

Multimodal Motion Prediction with Stacked Transformers

2021-03-22 · CVPR 2021 1 · Yicheng Liu, Jinghuai Zhang, Liangji Fang, Qinhong Jiang 외

Predicting multiple plausible future trajectories of the nearby vehicles is crucial for the safety of autonomous driving. Recent motion prediction approaches attempt to achieve such multimodal motion prediction by implic…

Autonomous DrivingDiversitymotion predictionPrediction

STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning

2025-12-15 · Jie Qin, Jiancheng Huang, Limeng Qiao, Lin Ma arxiv

Multimodal large language models (MLLMs) play a pivotal role in advancing the quest for general artificial intelligence. However, achieving unified target for multimodal understanding and generation remains challenging d…