paper-with-me

Papers

Improving Context Modelling in Multimodal Dialogue Generation

2018-10-20 · WS 2018 11 · Shubham Agarwal, Ondrej Dusek, Ioannis Konstas, Verena Rieser

In this work, we investigate the task of textual response generation in a multimodal task-oriented dialogue system. Our work is based on the recently released Multimodal Dialogue (MMD) dataset (Saha et al., 2017) in the fashion domain. We introduce a multimodal extension to the Hierarchical Recurrent Encoder-Decoder (HRED) model and show that this extension outperforms strong baselines in terms of text-based similarity metrics. We also showcase the shortcomings of current vision and language models by performing an error analysis on our system's output.

📄 PDF Abstract BibTeX arXiv:1810.11955

Code (1)

shubhamagarwal92/mmd 공식 구현 pytorch

Tasks

DecoderDialogue GenerationResponse Generation

Similar Papers 제목 키워드 기반

Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation

2024-05-16 · Bo Zhang, Hui Ma, Jian Ding, Jian Wang 외

Integrating multimodal knowledge into large language models (LLMs) represents a significant advancement in dialogue generation capabilities. However, the effective incorporation of such knowledge in zero-resource scenari…

Dialogue GenerationKnowledge Distillation

Multi-modal Generation via Cross-Modal In-Context Learning

2024-05-28 · Amandeep Kumar, Muzammal Naseer, Sanath Narayan, Rao Muhammad Anwer 외

In this work, we study the problem of generating novel images from complex multimodal prompt sequences. While existing methods achieve promising results for text-to-image generation, they often struggle to capture fine-g…

Image GenerationIn-Context LearningStory GenerationText to Image Generation+1

A Unified Framework for Slot based Response Generation in a Multimodal Dialogue System

2023-05-27 · Mauajama Firdaus, Avinash Madasu, Asif Ekbal

Natural Language Understanding (NLU) and Natural Language Generation (NLG) are the two critical components of every conversational system that handles the task of understanding the user by capturing the necessary informa…

DecoderNatural Language UnderstandingResponse GenerationText Generation

Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling

2024-12-30 · Min Zhang, Zilin Wang, Liyan Chen, KunHong Liu 외

Recent advances in AI-driven storytelling have enhanced video generation and story visualization. However, translating dialogue-centric scripts into coherent storyboards remains a significant challenge due to limited scr…

Retrieval-augmented GenerationStory VisualizationVideo Generation

BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation

2024-08-12 · Hee Suk Yoon, Eunseop Yoon, Joshua Tian Jin Tee, Kang Zhang 외

Multimodal Dialogue Response Generation (MDRG) is a recently proposed task where the model needs to generate responses in texts, images, or a blend of both based on the dialogue context. Due to the lack of a large-scale …

Response Generation