paper-with-me

Papers Visual Dialog

“Visual Dialog” 태그가 달린 논문 120편 · 필터 해제

FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

2026-07-30 · Bohan Hou, Haoqiang Lin, Xuemeng Song, Haokun Wen 외 arxiv

Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing d…

Image RetrievalVisual Dialog

Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments

2026-05-20 · Van Quang Nguyen arxiv

Advancements at the intersection of computer vision and natural language processing are crucial for applications like assistive tech, multimedia querying, and robotics. This dissertation proposes novel architectures to i…

Instruction FollowingImage CaptioningVisual Dialog

V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts

2025-03-03 · Adnen Abdessaied, Anna Rohrbach, Marcus Rohrbach, Andreas Bulling

We present V$^2$Dial - a novel expert-based model specifically geared towards simultaneously handling image and video input data for multimodal conversational tasks. Current multimodal models primarily focus on simpler t…

Contrastive LearningText RetrievalVideo-Text RetrievalVisual Dialog+1

V^2Dial: Unification of Video and Visual Dialog via Multimodal Experts

2025-01-01 · CVPR 2025 1 · Adnen Abdessaied, Anna Rohrbach, Marcus Rohrbach, Andreas Bulling

We present V2Dial - a novel expert-based model specifically geared towards simultaneously handling image and video input data for multimodal conversational tasks. Current multimodal models primarily focus on simpler …

Contrastive LearningText RetrievalVideo-Text RetrievalVisual Dialog+1

Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations

2024-08-13 · Wei Pang, Ruixue Duan, Jinfu Yang, Ning li

Visual Dialog (VD) is a task where an agent answers a series of image-related questions based on a multi-round dialog history. However, previous VD methods often treat the entire dialog history as a simple text input, di…

dialog state trackingDialogue State TrackingEntity AlignmentVisual Dialog

ICCV23 Visual-Dialog Emotion Explanation Challenge: SEU_309 Team Technical Report

2024-07-13 · Yixiao Yuan, Yingzhe Peng

The Visual-Dialog Based Emotion Explanation Generation Challenge focuses on generating emotion explanations through visual-dialog interactions in art discussions. Our approach combines state-of-the-art multi-modal models…

Explanation GenerationLanguage ModelingLanguage ModellingVisual Dialog

Hawk: Learning to Understand Open-World Video Anomalies

2024-05-27 · Jiaqi Tang, Hao Lu, Ruizheng Wu, Xiaogang Xu 외

Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial seman…

Anomaly DetectionQuestion AnsweringVideo Anomaly DetectionVideo Description+2

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

2024-03-27 · Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong 외

In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic visual dialog and reasoning, a performa…

Image ClassificationImage ComprehensionReferring Expression ComprehensionReferring expression generation+2

FlexCap: Describe Anything in Images in Controllable Detail

2024-03-18 · Debidatta Dwibedi, Vidhi Jain, Jonathan Tompson, Andrew Zisserman 외

We introduce FlexCap, a vision-language model that generates region-specific descriptions of varying lengths. FlexCap is trained to produce length-conditioned captions for input boxes, enabling control over information d…

AttributeDense CaptioningLanguage ModelingLanguage Modelling+8

$\mathbb{VD}$-$\mathbb{GR}$: Boosting $\mathbb{V}$isual $\mathbb{D}$ialog with Cascaded Spatial-Temporal Multi-Modal $\mathbb{GR}$aphs

2023-10-25 · Adnen Abdessaied, Lei Shi, Andreas Bulling

We propose $\mathbb{VD}$-$\mathbb{GR}$ - a novel visual dialog model that combines pre-trained language models (LMs) with graph neural networks (GNNs). Prior works mainly focused on one class of models at the expense of …

Visual Dialog

Collecting Visually-Grounded Dialogue with A Game Of Sorts

2023-09-10 · LREC 2022 6 · Bram Willemsen, Dmytro Kalpakchi, Gabriel Skantze

An idealized, though simplistic, view of the referring expression production and grounding process in (situated) dialogue assumes that a speaker must merely appropriately specify their expression so that the target refer…

Coreference ResolutionImage RetrievalReferring ExpressionReferring Expression Comprehension+4

Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations

2023-08-30 · Kilichbek Haydarov, Xiaoqian Shen, Avinash Madasu, Mahmoud Salem 외

We introduce Affective Visual Dialog, an emotion explanation and reasoning task as a testbed for research on understanding the formation of emotions in visually grounded conversations. The task involves three skills: (1)…

Explanation GenerationQuestion AnsweringVisual Dialog

PaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional Experts

2023-05-24 · Yunshui Li, Binyuan Hui, Zhichao Yin, Min Yang 외

Perceiving multi-modal information and fulfilling dialogues with humans is a long-term goal of artificial intelligence. Pre-training is commonly regarded as an effective approach for multi-modal dialogue. However, due to…

Dialogue State TrackingImage RetrievalMultimodal Intent RecognitionResponse Generation+2

Unified Multimodal Model with Unlikelihood Training for Visual Dialog

2022-11-23 · ZiHao Wang, Junli Wang, Changjun Jiang

The task of visual dialog requires a multimodal chatbot to answer sequential questions from humans about image content. Prior work performs the standard likelihood training for answer generation on the positive instances…

Answer GenerationChatbotLanguage ModelingLanguage Modelling+3

A survey on knowledge-enhanced multimodal learning

2022-11-19 · Maria Lymperaiou, Giorgos Stamou

Multimodal learning has been a field of increasing interest, aiming to combine various modalities in a single joint representation. Especially in the area of visiolinguistic (VL) learning multiple models and techniques h…

Conditional Image GenerationFactual Visual Question AnsweringFairnessImage Captioning+13

Knowledge Transfer with Visual Prompt in multi-modal Dialogue Understanding and Generation

2022-10-01 · TU (COLING) 2022 10 · Minjun Zhu, Yixuan Weng, Bin Li, Shizhu He 외

Visual Dialogue (VD) task has recently received increasing attention in AI research. Visual Dialog aims to generate multi-round, interactive responses based on the dialog history and image content. Existing textual dialo…

Dialogue UnderstandingKnowledge DistillationTransfer LearningVisual Dialog

LAVIS: A Library for Language-Vision Intelligence

2022-09-15 · Dongxu Li, Junnan Li, Hung Le, Guangsen Wang 외

We introduce LAVIS, an open-source deep learning library for LAnguage-VISion research and applications. LAVIS aims to serve as a one-stop comprehensive library that brings recent advancements in the language-vision field…

BenchmarkingImage CaptioningImage RetrievalMultimodal Deep Learning+7

Video Dialog as Conversation about Objects Living in Space-Time

2022-07-08 · Hoang-Anh Pham, Thao Minh Le, Vuong Le, Tu Minh Phuong 외

It would be a technological feat to be able to create a system that can hold a meaningful conversation with humans about what they watch. A setup toward that goal is presented as a video dialog task, where the system is …

ObjectRelational ReasoningRetrievalVideo Question Answering+1

Adversarial Robustness of Visual Dialog

2022-07-06 · Lu Yu, Verena Rieser

Adversarial robustness evaluates the worst-case performance scenario of a machine learning model to ensure its safety and reliability. This study is the first to investigate the robustness of visually grounded dialog mod…

Adversarial RobustnessVisual Dialog

ENRICH4ALL: A First Luxembourgish BERT Model for a Multilingual Chatbot

2022-06-01 · SIGUL (LREC) 2022 6 · Dimitra Anastasiou

Machine Translation (MT)-empowered chatbots are not established yet, however, we see an amazing future breaking language barriers and enabling conversation in multiple languages without time-consuming language model buil…

ChatbotLanguage ModelingLanguage ModellingMachine Translation+2
1–20 / 120 다음 →