paper-with-me

홈 › Papers

Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue

2026-04-22 · Biswesh Mohapatra, Giovanni Duca, Laurent Romary, Justine Cassell arxiv

Situated dialogue requires speakers to maintain a reliable representation of shared context rather than reasoning only over isolated utterances. Current conversational agents often struggle with this requirement, especially when the common ground must be preserved beyond the immediate context window. In such settings, fine-grained distinctions are frequently compressed into purely textual representations, leading to a critical failure mode we call \emph{representational blur}, in which similar but distinct entities collapse into interchangeable descriptions. This semantic flattening creates an illusion of grounding, where agents appear locally coherent but fail to track shared context persistently over time. Inspired by the role of mental imagery in human reasoning, and based on the increased availability of multimodal models, we explore whether conversational agents can be given an analogous ability to construct some depictive intermediate representations during dialogue to address these limitations. Thus, we introduce an active visual scaffolding framework that incrementally converts dialogue state into a persistent visual history that can later be retrieved for grounded response generation. Evaluation on the IndiRef benchmark shows that incremental externalization itself improves over full-dialog reasoning, while visual scaffolding provides additional gains by reducing representational blur and enforcing concrete scene commitments. At the same time, textual representations remain advantageous for non-depictable information, and a hybrid multimodal setting yields the best overall performance. Together, these findings suggest that conversational agents benefit from an explicitly multimodal representation of common ground that integrates depictive and propositional information.

📄 PDF Abstract BibTeX arXiv:2604.21144

Code (0)

등록된 구현이 없습니다.

Tasks

Response Generation

Similar Papers 제목 키워드 기반

Towards Asynchronous Motor Imagery-Based Brain-Computer Interfaces: a joint training scheme using deep learning

2018-08-31 · Patcharin Cheng, Phairot Autthasan, Boriwat Pijarana, Ekapol Chuangsuwanich 외

In this paper, the deep learning (DL) approach is applied to a joint training scheme for asynchronous motor imagery-based Brain-Computer Interface (BCI). The proposed DL approach is a cascade of one-dimensional convoluti…

Brain Computer InterfaceEEGElectroencephalogram (EEG)Motor Imagery

Model-based Influence Diagrams for Machine Vision

2013-03-27 · Tod S. Levitt, John Mark Agosta, Thomas O. Binford

We show an approach to automated control of machine vision systems based on incremental creation and evaluation of a particular family of influence diagrams that represent hypotheses of imagery interpretation and possibl…

Bayesian Inferencemodel

Assessing thermal imagery integration into object detection methods on ground-based and air-based collection platforms

2022-12-23 · James Gallagher, Edward Oughton

Object detection models commonly deployed on uncrewed aerial systems (UAS) focus on identifying objects in the visible spectrum using Red-Green-Blue (RGB) imagery. However, there is growing interest in fusing RGB with th…

Objectobject-detectionObject Detection

Population Mapping in Informal Settlements with High-Resolution Satellite Imagery and Equitable Ground-Truth

2020-09-17 · Konstantin Klemmer, Godwin Yeboah, João Porto de Albuquerque, Stephen A Jarvis

We propose a generalizable framework for the population estimation of dense, informal settlements in low-income urban areas--so called 'slums'--using high-resolution satellite imagery. Precise population estimates are a …

BIG-bench Machine LearningPopulation Mapping

Generating Synthetic Multispectral Satellite Imagery from Sentinel-2

2020-12-05 · Tharun Mohandoss, Aditya Kulkarni, Daniel Northrup, Ernest Mwebaze 외

Multi-spectral satellite imagery provides valuable data at global scale for many environmental and socio-economic applications. Building supervised machine learning models based on these imagery, however, may require gro…

BIG-bench Machine LearningData Augmentation