paper-with-me

홈 › Papers

Distraction-free Embeddings for Robust VQA

2023-08-31 · Atharvan Dogra, Deeksha Varshney, Ashwin Kalyan, Ameet Deshpande, Neeraj Kumar

The generation of effective latent representations and their subsequent refinement to incorporate precise information is an essential prerequisite for Vision-Language Understanding (VLU) tasks such as Video Question Answering (VQA). However, most existing methods for VLU focus on sparsely sampling or fine-graining the input information (e.g., sampling a sparse set of frames or text tokens), or adding external knowledge. We present a novel "DRAX: Distraction Removal and Attended Cross-Alignment" method to rid our cross-modal representations of distractors in the latent space. We do not exclusively confine the perception of any input information from various modalities but instead use an attention-guided distraction removal method to increase focus on task-relevant information in latent embeddings. DRAX also ensures semantic alignment of embeddings during cross-modal fusions. We evaluate our approach on a challenging benchmark (SUTD-TrafficQA dataset), testing the framework's abilities for feature and event queries, temporal relation understanding, forecasting, hypothesis, and causal analysis through extensive experiments.

📄 PDF Abstract BibTeX arXiv:2309.00133

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVideo Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Predicting Visual Attention and Distraction During Visual Search Using Convolutional Neural Networks

2022-10-27 · Manoosh Samiei, James J. Clark

Most studies in computational modeling of visual attention encompass task-free observation of images. Free-viewing saliency considers limited scenarios of daily life. Most visual activities are goal-oriented and demand a…

DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical Representations

2021-10-27 · Fei Deng, Ingook Jang, Sungjin Ahn

Top-performing Model-Based Reinforcement Learning (MBRL) agents, such as Dreamer, learn the world model by reconstructing the image observations. Hence, they often fail to discard task-irrelevant details and struggle to …

Model-based Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Classification of Distraction Levels Using Hybrid Deep Neural Networks From EEG Signals

2022-12-13 · Dae-Hyeok Lee, Sung-Jin Kim, Yeon-Woo Choi

Non-invasive brain-computer interface technology has been developed for detecting human mental states with high performances. Detection of the pilots' mental states is particularly critical because their abnormal mental …

Autonomous DrivingBrain Computer InterfaceEEGElectroencephalogram (EEG)

Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective

2025-11-20 · Jiahao Li, Yang Lu, Yachao Zhang, Yong Xie 외 arxiv

Open-vocabulary semantic segmentation (OVSS) employs pixel-level vision-language alignment to associate category-related prompts with corresponding pixels. A key challenge is enhancing the multimodal dense prediction cap…

Semantic Segmentation

A Geometric Perspective on Self-Supervised Policy Adaptation

2020-11-14 · Cristian Bodnar, Karol Hausman, Gabriel Dulac-Arnold, Rico Jonschkowski

One of the most challenging aspects of real-world reinforcement learning (RL) is the multitude of unpredictable and ever-changing distractions that could divert an agent from what was tasked to do in its training environ…

Reinforcement Learning (RL)