Variational Disentangled Attention for Regularized Visual Dialog
One of the most important challenges in a visual dialog is to effectively extract the information from a given image and its historical conversation which are related to the current question. Many studies adopt the soft attention mechanism in different information sources due to its simplicity and ease of optimization. However, some of visual dialogs are observed in a single round. This implies that there is no substantial correlation between individual rounds of questions and answers. This paper presents a unified approach to disentangled attention to deal with context-free visual dialogs. The question is disentangled in latent representation. In particular, an informative regularization is imposed to strengthen the dependence between vision and language by pretraining on the visual question answering before transferring to visual dialog. Importantly, a novel variational attention mechanism is developed and implemented by a local reparameterization trick which carries out a discrete attention to identify the relevant conversations in a visual dialog. A set of experiments are evaluated to illustrate the merits of the proposed attention and regularization schemes for context-free visual dialogs.
Code (0)
등록된 구현이 없습니다.
Tasks
Question AnsweringVisual DialogVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Semantically Conditioned Dialog Response Generation via Hierarchical Disentangled Self-Attention
Semantically controlled neural response generation on limited-domain has achieved great performance. However, moving towards multi-domain large-scale scenarios are shown to be difficult because the possible combinations …
Data-to-Text GenerationInductive BiasResponse GenerationDualVAE: Dual Disentangled Variational AutoEncoder for Recommendation
Learning precise representations of users and items to fit observed interaction data is the fundamental task of collaborative filtering. Existing studies usually infer entangled representations to fit such interaction da…
Collaborative FilteringDisentanglementRepresentation LearningVariational InferenceMoToRec: Sparse-Regularized Multimodal Tokenization for Cold-Start Recommendation
Graph neural networks (GNNs) have revolutionized recommender systems by effectively modeling complex user-item interactions, yet data sparsity and the item cold-start problem significantly impair performance, particularl…
Multimodal RecommendationAttri-VAE: attribute-based interpretable representations of medical images with variational autoencoders
Deep learning (DL) methods where interpretability is intrinsically considered as part of the model are required to better understand the relationship of clinical and imaging-based attributes with DL outcomes, thus facili…
AttributeDisentanglementDisentangling Generative Factors of Physical Fields Using Variational Autoencoders
The ability to extract generative parameters from high-dimensional fields of data in an unsupervised manner is a highly desirable yet unrealized goal in computational physics. This work explores the use of variational au…
Dimensionality ReductionDisentanglement