Attending to Visual Differences for Situated Language Generation in Changing Scenes
We investigate the problem of generating utterances from pairs of images showing a before and an after state of a change in a visual scene. We present a transformer model with difference attention heads that learns to attend to visual changes in consecutive images via a difference key. We test our approach in instruction generation, change captioning, and difference spotting and compare these tasks in terms of their linguistic phenomena and reasoning abilities. Our model outperforms the state-of-the-art for instruction generation on the BLOCKS and difference spotting on the Spot-the-diff dataset and generates accurate referential and compositional spatial expressions. Finally, we identify linguistic phenomena that pose challenges for generation in changing scenes.
Code (0)
등록된 구현이 없습니다.
Tasks
Text GenerationSimilar Papers 제목 키워드 기반
Incremental Generation of Visually Grounded Language in Situated Dialogue (demonstration system)
UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Text is ubiquitous in our visual world, conveying crucial information, such as in documents, websites, and everyday photographs. In this work, we propose UReader, a first exploration of universal OCR-free visually-situat…
DecoderLanguage ModelingLanguage ModellingLarge Language Model+2MindDial: Belief Dynamics Tracking with Theory-of-Mind Modeling for Situated Neural Dialogue Generation
Humans talk in daily conversations while aligning and negotiating the expressed meanings or common ground. Despite the impressive conversational abilities of the large generative language models, they do not consider the…
Dialogue GenerationTheory of Mind ModelingOmniParser: A Unified Framework for Text Spotting Key Information Extraction and Table Recognition
Recently visually-situated text parsing (VsTP) has experienced notable advancements driven by the increasing demand for automated document understanding and the emergence of Generative Large Language Models (LLMs) ca…
Decoderdocument understandingKey Information ExtractionTable Recognition+2OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
Recently, visually-situated text parsing (VsTP) has experienced notable advancements, driven by the increasing demand for automated document understanding and the emergence of Generative Large Language Models (LLMs) capa…
Decoderdocument understandingKey Information ExtractionTable Recognition+2