Papers Visual Dialog
“Visual Dialog” 태그가 달린 논문 120편 · 필터 해제
FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing d…
Image RetrievalVisual DialogMachine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments
Advancements at the intersection of computer vision and natural language processing are crucial for applications like assistive tech, multimedia querying, and robotics. This dissertation proposes novel architectures to i…
Instruction FollowingImage CaptioningVisual DialogV$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
We present V$^2$Dial - a novel expert-based model specifically geared towards simultaneously handling image and video input data for multimodal conversational tasks. Current multimodal models primarily focus on simpler t…
Contrastive LearningText RetrievalVideo-Text RetrievalVisual Dialog+1V^2Dial: Unification of Video and Visual Dialog via Multimodal Experts
We present V2Dial - a novel expert-based model specifically geared towards simultaneously handling image and video input data for multimodal conversational tasks. Current multimodal models primarily focus on simpler …
Contrastive LearningText RetrievalVideo-Text RetrievalVisual Dialog+1Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations
Visual Dialog (VD) is a task where an agent answers a series of image-related questions based on a multi-round dialog history. However, previous VD methods often treat the entire dialog history as a simple text input, di…
dialog state trackingDialogue State TrackingEntity AlignmentVisual DialogICCV23 Visual-Dialog Emotion Explanation Challenge: SEU_309 Team Technical Report
The Visual-Dialog Based Emotion Explanation Generation Challenge focuses on generating emotion explanations through visual-dialog interactions in art discussions. Our approach combines state-of-the-art multi-modal models…
Explanation GenerationLanguage ModelingLanguage ModellingVisual DialogHawk: Learning to Understand Open-World Video Anomalies
Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial seman…
Anomaly DetectionQuestion AnsweringVideo Anomaly DetectionVideo Description+2Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic visual dialog and reasoning, a performa…
Image ClassificationImage ComprehensionReferring Expression ComprehensionReferring expression generation+2FlexCap: Describe Anything in Images in Controllable Detail
We introduce FlexCap, a vision-language model that generates region-specific descriptions of varying lengths. FlexCap is trained to produce length-conditioned captions for input boxes, enabling control over information d…
AttributeDense CaptioningLanguage ModelingLanguage Modelling+8$\mathbb{VD}$-$\mathbb{GR}$: Boosting $\mathbb{V}$isual $\mathbb{D}$ialog with Cascaded Spatial-Temporal Multi-Modal $\mathbb{GR}$aphs
We propose $\mathbb{VD}$-$\mathbb{GR}$ - a novel visual dialog model that combines pre-trained language models (LMs) with graph neural networks (GNNs). Prior works mainly focused on one class of models at the expense of …
Visual DialogCollecting Visually-Grounded Dialogue with A Game Of Sorts
An idealized, though simplistic, view of the referring expression production and grounding process in (situated) dialogue assumes that a speaker must merely appropriately specify their expression so that the target refer…
Coreference ResolutionImage RetrievalReferring ExpressionReferring Expression Comprehension+4Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations
We introduce Affective Visual Dialog, an emotion explanation and reasoning task as a testbed for research on understanding the formation of emotions in visually grounded conversations. The task involves three skills: (1)…
Explanation GenerationQuestion AnsweringVisual DialogPaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional Experts
Perceiving multi-modal information and fulfilling dialogues with humans is a long-term goal of artificial intelligence. Pre-training is commonly regarded as an effective approach for multi-modal dialogue. However, due to…
Dialogue State TrackingImage RetrievalMultimodal Intent RecognitionResponse Generation+2Unified Multimodal Model with Unlikelihood Training for Visual Dialog
The task of visual dialog requires a multimodal chatbot to answer sequential questions from humans about image content. Prior work performs the standard likelihood training for answer generation on the positive instances…
Answer GenerationChatbotLanguage ModelingLanguage Modelling+3A survey on knowledge-enhanced multimodal learning
Multimodal learning has been a field of increasing interest, aiming to combine various modalities in a single joint representation. Especially in the area of visiolinguistic (VL) learning multiple models and techniques h…
Conditional Image GenerationFactual Visual Question AnsweringFairnessImage Captioning+13Knowledge Transfer with Visual Prompt in multi-modal Dialogue Understanding and Generation
Visual Dialogue (VD) task has recently received increasing attention in AI research. Visual Dialog aims to generate multi-round, interactive responses based on the dialog history and image content. Existing textual dialo…
Dialogue UnderstandingKnowledge DistillationTransfer LearningVisual DialogLAVIS: A Library for Language-Vision Intelligence
We introduce LAVIS, an open-source deep learning library for LAnguage-VISion research and applications. LAVIS aims to serve as a one-stop comprehensive library that brings recent advancements in the language-vision field…
BenchmarkingImage CaptioningImage RetrievalMultimodal Deep Learning+7Video Dialog as Conversation about Objects Living in Space-Time
It would be a technological feat to be able to create a system that can hold a meaningful conversation with humans about what they watch. A setup toward that goal is presented as a video dialog task, where the system is …
ObjectRelational ReasoningRetrievalVideo Question Answering+1Adversarial Robustness of Visual Dialog
Adversarial robustness evaluates the worst-case performance scenario of a machine learning model to ensure its safety and reliability. This study is the first to investigate the robustness of visually grounded dialog mod…
Adversarial RobustnessVisual DialogENRICH4ALL: A First Luxembourgish BERT Model for a Multilingual Chatbot
Machine Translation (MT)-empowered chatbots are not established yet, however, we see an amazing future breaking language barriers and enabling conversation in multiple languages without time-consuming language model buil…
ChatbotLanguage ModelingLanguage ModellingMachine Translation+2