paper-with-me

홈 › Papers

DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue

2019-11-17 · Xiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang, Xingxing Zhang, Yue Hu, Qi Wu

Different from Visual Question Answering task that requires to answer only one question about an image, Visual Dialogue involves multiple questions which cover a broad range of visual content that could be related to any objects, relationships or semantics. The key challenge in Visual Dialogue task is thus to learn a more comprehensive and semantic-rich image representation which may have adaptive attentions on the image for variant questions. In this research, we propose a novel model to depict an image from both visual and semantic perspectives. Specifically, the visual view helps capture the appearance-level information, including objects and their relationships, while the semantic view enables the agent to understand high-level visual semantics from the whole image to the local regions. Futhermore, on top of such multi-view image features, we propose a feature selection framework which is able to adaptively capture question-relevant information hierarchically in fine-grained level. The proposed method achieved state-of-the-art results on benchmark Visual Dialogue datasets. More importantly, we can tell which modality (visual or semantic) has more contribution in answering the current question by visualizing the gate values. It gives us insights in understanding of human cognition in Visual Dialogue.

📄 PDF Abstract BibTeX arXiv:1911.07251

Code (1)

JXZe/DualVD 공식 구현 pytorch

Tasks

feature selectionQuestion AnsweringVisual DialogVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Feature Selection Feature selection, also known as variable selection, attribute selection or variable subset selection, is the process of selecting a subset of relevant features (variables,…

Similar Papers 제목 키워드 기반

Dual reparametrized Variational Generative Model for Time-Series Forecasting

2022-03-11 · Ziang Chen

This paper propose DualVDT, a generative model for Time-series forecasting. Introduced dual reparametrized variational mechanisms on variational autoencoder (VAE) to tighter the evidence lower bound (ELBO) of the model, …

DenoisingTime SeriesTime Series AnalysisTime Series Forecasting

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs

2026-04-13 · Haicheng Wang, Yuan Liu, Yikun Liu, Zhemeng Yu 외 arxiv

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the rapid growth of visual token sequences--especially in long-video and s…

Visual Reasoning

Lance: Unified Multimodal Modeling by Multi-Task Synergy

2026-05-18 · Fengyi Fu, Mengqi Huang, Shaojin Wu, Yunsheng Jiang 외 arxiv

We present Lance, a lightweight native unified model supporting multimodal understanding, generation, and editing for both images and videos. Rather than relying on model capacity scaling or text-image-dominant designs, …

Video Generation

Adaptive Dual-Path Framework for Covert Semantic Communication

2026-05-05 · Xi Yu, Weicai Li, Lin Yin, Tiejun Lv arxiv

This paper proposes a novel adaptive dual-path framework for covert semantic communication (SemCom), which integrates covert information transmission with task-oriented semantic coding. Unlike conventional covert communi…

Semantic Communication

PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference

2026-08-27 · Junjie Liu, Shengyuan Ye, Xu Chen arxiv

Vision-Language Models (VLMs) demonstrate exceptional visual reasoning capabilities, yet their inference costs escalate rapidly with the proliferation of visual tokens. Existing visual token pruning methods exhibit two f…

Visual Reasoning