paper-with-me

홈 › Papers

ReSee: Responding through Seeing Fine-grained Visual Knowledge in Open-domain Dialogue

2023-05-23 · Haoqin Tu, Yitong Li, Fei Mi, Zhongliang Yang

Incorporating visual knowledge into text-only dialogue systems has become a potential direction to imitate the way humans think, imagine, and communicate. However, existing multimodal dialogue systems are either confined by the scale and quality of available datasets or the coarse concept of visual knowledge. To address these issues, we provide a new paradigm of constructing multimodal dialogues as well as two datasets extended from text-only dialogues under such paradigm (ReSee-WoW, ReSee-DD). We propose to explicitly split the visual knowledge into finer granularity (`turn-level'' and `entity-level''). To further boost the accuracy and diversity of augmented visual information, we retrieve them from the Internet or a large image dataset. To demonstrate the superiority and universality of the provided visual knowledge, we propose a simple but effective framework ReSee to add visual representation into vanilla dialogue models by modality concatenations. We also conduct extensive experiments and ablations w.r.t. different model configurations and visual knowledge settings. Empirical, encouraging results not only demonstrate the effectiveness of introducing visual knowledge at both entity and turn level but also verify the proposed model ReSee outperforms several state-of-the-art methods on automatic and human evaluations. By leveraging text and vision knowledge, ReSee can produce informative responses with real-world visual concepts. Our code is available at https://github.com/ImKeTT/ReSee.

📄 PDF Abstract BibTeX arXiv:2305.13602

Code (1)

imkett/resee 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Foreseeing Brain Graph Evolution Over Time Using Deep Adversarial Network Normalizer

2020-09-23 · Zeynep Gurler, Ahmed Nebli, Islem Rekik

Foreseeing the brain evolution as a complex highly inter-connected system, widely modeled as a graph, is crucial for mapping dynamic interactions between different anatomical regions of interest (ROIs) in health and dise…

Generative Adversarial Network

Decoding Large Language Diffusion Models with Foreseeing Movement

2025-12-03 · Yichuan Mo, Quan Chen, Mingjie Li, Zeming Wei 외 arxiv

Large Language Diffusion Models (LLDMs) benefit from a flexible decoding mechanism that enables parallelized inference and controllable generations over autoregressive models. Yet such flexibility introduces a critical c…

Plan To Predict: Learning an Uncertainty-Foreseeing Model for Model-Based Reinforcement Learning

2023-01-20 · Zifan Wu, Chao Yu, Chen Chen, Jianye Hao 외

In Model-based Reinforcement Learning (MBRL), model learning is critical since an inaccurate model can bias policy learning via generating misleading samples. However, learning an accurate model can be difficult since th…

Decision MakingmodelModel-based Reinforcement LearningSequential Decision Making

Fine-Grained ImageNet Classification in the Wild

2023-03-04 · Maria Lymperaiou, Konstantinos Thomas, Giorgos Stamou

Image classification has been one of the most popular tasks in Deep Learning, seeing an abundance of impressive implementations each year. However, there is a lot of criticism tied to promoting complex architectures that…

Classificationimage-classificationImage Classification

From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation

2025-10-01 · Fan Yang, Zhiyang Chen, Yousong Zhu, Xin Li 외 arxiv

Current video generation models produce physically inconsistent motion that violates real-world dynamics. We propose TrajVLM-Gen, a two-stage framework for physics-aware image-to-video generation. First, we employ a Visi…

Trajectory ForecastingTrajectory PredictionVideo Generation