paper-with-me

홈 › Papers

Improving Cross-Modal Understanding in Visual Dialog via Contrastive Learning

2022-04-15 · Feilong Chen, Xiuyi Chen, Shuang Xu, Bo Xu

Visual Dialog is a challenging vision-language task since the visual dialog agent needs to answer a series of questions after reasoning over both the image content and dialog history. Though existing methods try to deal with the cross-modal understanding in visual dialog, they are still not enough in ranking candidate answers based on their understanding of visual and textual contexts. In this paper, we analyze the cross-modal understanding in visual dialog based on the vision-language pre-training model VD-BERT and propose a novel approach to improve the cross-modal understanding for visual dialog, named ICMU. ICMU enhances cross-modal understanding by distinguishing different pulled inputs (i.e. pulled images, questions or answers) based on four-way contrastive learning. In addition, ICMU exploits the single-turn visual question answering to enhance the visual dialog model's cross-modal understanding to handle a multi-turn visually-grounded conversation. Experiments show that the proposed approach improves the visual dialog model's cross-modal understanding and brings satisfactory gain to the VisDial dataset.

📄 PDF Abstract BibTeX arXiv:2204.07302

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningQuestion AnsweringVisual DialogVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

$C^3$: Compositional Counterfactual Contrastive Learning for Video-grounded Dialogues

2021-06-16 · Hung Le, Nancy F. Chen, Steven C. H. Hoi

Video-grounded dialogue systems aim to integrate video understanding and dialogue understanding to generate responses that are relevant to both the dialogue and video context. Most existing approaches employ deep learnin…

Contrastive LearningcounterfactualDialogue UnderstandingMultimodal Reasoning+1

Knowledge Transfer with Visual Prompt in multi-modal Dialogue Understanding and Generation

2022-10-01 · TU (COLING) 2022 10 · Minjun Zhu, Yixuan Weng, Bin Li, Shizhu He 외

Visual Dialogue (VD) task has recently received increasing attention in AI research. Visual Dialog aims to generate multi-round, interactive responses based on the dialog history and image content. Existing textual dialo…

Dialogue UnderstandingKnowledge DistillationTransfer LearningVisual Dialog

SPACE-2: Tree-Structured Semi-Supervised Contrastive Pre-training for Task-Oriented Dialog Understanding

2022-09-14 · COLING 2022 10 · Wanwei He, Yinpei Dai, Binyuan Hui, Min Yang 외

Pre-training methods with contrastive learning objectives have shown remarkable success in dialog understanding tasks. However, current contrastive learning solely considers the self-augmented dialog samples as positive …

Contrastive LearningSTS

Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling

2024-12-30 · Min Zhang, Zilin Wang, Liyan Chen, KunHong Liu 외

Recent advances in AI-driven storytelling have enhanced video generation and story visualization. However, translating dialogue-centric scripts into coherent storyboards remains a significant challenge due to limited scr…

Retrieval-augmented GenerationStory VisualizationVideo Generation

ZRIGF: An Innovative Multimodal Framework for Zero-Resource Image-Grounded Dialogue Generation

2023-08-01 · Bo Zhang, Jian Wang, Hui Ma, Bo Xu 외

Image-grounded dialogue systems benefit greatly from integrating visual information, resulting in high-quality response generation. However, current models struggle to effectively utilize such information in zero-resourc…

Dialogue GenerationResponse Generation