paper-with-me

Papers

Variational Disentangled Attention for Regularized Visual Dialog

2021-09-29 · Jen-Tzung Chien, Hsiu-Wei Tien

One of the most important challenges in a visual dialog is to effectively extract the information from a given image and its historical conversation which are related to the current question. Many studies adopt the soft attention mechanism in different information sources due to its simplicity and ease of optimization. However, some of visual dialogs are observed in a single round. This implies that there is no substantial correlation between individual rounds of questions and answers. This paper presents a unified approach to disentangled attention to deal with context-free visual dialogs. The question is disentangled in latent representation. In particular, an informative regularization is imposed to strengthen the dependence between vision and language by pretraining on the visual question answering before transferring to visual dialog. Importantly, a novel variational attention mechanism is developed and implemented by a local reparameterization trick which carries out a discrete attention to identify the relevant conversations in a visual dialog. A set of experiments are evaluated to illustrate the merits of the proposed attention and regularization schemes for context-free visual dialogs.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual DialogVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Semantically Conditioned Dialog Response Generation via Hierarchical Disentangled Self-Attention

2019-05-30 · ACL 2019 7 · Wenhu Chen, Jianshu Chen, Pengda Qin, Xifeng Yan 외

Semantically controlled neural response generation on limited-domain has achieved great performance. However, moving towards multi-domain large-scale scenarios are shown to be difficult because the possible combinations …

Data-to-Text GenerationInductive BiasResponse Generation

DualVAE: Dual Disentangled Variational AutoEncoder for Recommendation

2024-01-10 · Zhiqiang Guo, GuoHui Li, Jianjun Li, Chaoyang Wang 외

Learning precise representations of users and items to fit observed interaction data is the fundamental task of collaborative filtering. Existing studies usually infer entangled representations to fit such interaction da…

Collaborative FilteringDisentanglementRepresentation LearningVariational Inference

MoToRec: Sparse-Regularized Multimodal Tokenization for Cold-Start Recommendation

2026-02-11 · Jialin Liu, Zhaorui Zhang, Ray C. C. Cheung arxiv

Graph neural networks (GNNs) have revolutionized recommender systems by effectively modeling complex user-item interactions, yet data sparsity and the item cold-start problem significantly impair performance, particularl…

Multimodal Recommendation

Attri-VAE: attribute-based interpretable representations of medical images with variational autoencoders

2022-03-20 · Irem Cetin, Maialen Stephens, Oscar Camara, Miguel Angel Gonzalez Ballester

Deep learning (DL) methods where interpretability is intrinsically considered as part of the model are required to better understand the relationship of clinical and imaging-based attributes with DL outcomes, thus facili…

AttributeDisentanglement

Disentangling Generative Factors of Physical Fields Using Variational Autoencoders

2021-09-15 · Christian Jacobsen, Karthik Duraisamy

The ability to extract generative parameters from high-dimensional fields of data in an unsupervised manner is a highly desirable yet unrealized goal in computational physics. This work explores the use of variational au…

Dimensionality ReductionDisentanglement