paper-with-me

Papers

Multimodal grid features and cell pointers for Scene Text Visual Question Answering

2020-06-01 · Lluís Gómez, Ali Furkan Biten, Rubèn Tito, Andrés Mafla, Marçal Rusiñol, Ernest Valveny, Dimosthenis Karatzas

This paper presents a new model for the task of scene text visual question answering, in which questions about a given image can only be answered by reading and understanding scene text that is present in it. The proposed model is based on an attention mechanism that attends to multi-modal features conditioned to the question, allowing it to reason jointly about the textual and visual modalities in the scene. The output weights of this attention module over the grid of multi-modal spatial features are interpreted as the probability that a certain spatial location of the image contains the answer text the to the given question. Our experiments demonstrate competitive performance in two standard datasets. Furthermore, this paper provides a novel analysis of the ST-VQA dataset based on a human performance study.

📄 PDF Abstract BibTeX arXiv:2006.00923

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Neurally-plausible radial basis kernels using distributed Fourier embeddings

2026-05-08 · Jakeb Chouinard arxiv

Coherent, continuous spatial representations are critical for synthesizing physical and perceptual phenomena into a single representational space. Radial basis kernels provide a path forward for this type of distributed …

Scene-LSTM: A Model for Human Trajectory Prediction

2018-08-12 · Huynh Manh, Gita Alaghband

We develop a human movement trajectory prediction system that incorporates the scene information (Scene-LSTM) as well as human movement trajectories (Pedestrian movement LSTM) in the prediction process within static crow…

modelPredictionTrajectory Prediction

Trajectory Prediction by Coupling Scene-LSTM with Human Movement LSTM

2019-08-23 · Manh Huynh, Gita Alaghband

We develop a novel human trajectory prediction system that incorporates the scene information (Scene-LSTM) as well as individual pedestrian movement (Pedestrian-LSTM) trained simultaneously within static crowded scenes. …

Trajectory Prediction

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

2026-07-21 · Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch, Srinath Sridhar 외 hf

Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved throug…

Image Generation

Long-Term Occupancy Grid Prediction Using Recurrent Neural Networks

2018-09-11 · Marcel Schreiber, Stefan Hoermann, Klaus Dietmayer

We tackle the long-term prediction of scene evolution in a complex downtown scenario for automated driving based on Lidar grid fusion and recurrent neural networks (RNNs). A bird's eye view of the scene, including occupa…

Prediction