paper-with-me

Papers

L2C: Describing Visual Differences Needs Semantic Understanding of Individuals

2021-02-03 · EACL 2021 2 · An Yan, Xin Eric Wang, Tsu-Jui Fu, William Yang Wang

Recent advances in language and vision push forward the research of captioning a single image to describing visual differences between image pairs. Suppose there are two images, I_1 and I_2, and the task is to generate a description W_{1,2} comparing them, existing methods directly model { I_1, I_2 } -> W_{1,2} mapping without the semantic understanding of individuals. In this paper, we introduce a Learning-to-Compare (L2C) model, which learns to understand the semantic structures of these two images and compare them while learning to describe each one. We demonstrate that L2C benefits from a comparison between explicit semantic representations and single-image captions, and generalizes better on the new testing image pairs. It outperforms the baseline on both automatic evaluation and human evaluation for the Birds-to-Words dataset.

📄 PDF Abstract BibTeX arXiv:2102.01860

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

FutureVision: A methodology for the investigation of future cognition

2025-02-03 · Tiago Timponi Torrent, Mark Turner, Nicolás Hinrichs, Frederico Belcavello 외

This paper presents a methodology combining multimodal semantic analysis with an eye-tracking experimental protocol to investigate the cognitive effort involved in understanding the communication of future scenarios. To …

Bridging Visual Perception with Contextual Semantics for Understanding Robot Manipulation Tasks

2019-09-16 · Chen Jiang, Martin Jagersand

Understanding manipulation scenarios allows intelligent robots to plan for appropriate actions to complete a manipulation task successfully. It is essential for intelligent robots to semantically interpret manipulation k…

AttributeCommon Sense ReasoningKnowledge GraphsLanguage Modeling+2

Audio Difference Captioning Utilizing Similarity-Discrepancy Disentanglement

2023-08-23 · Daiki Takeuchi, Yasunori Ohishi, Daisuke Niizumi, Noboru Harada 외

We proposed Audio Difference Captioning (ADC) as a new extension task of audio captioning for describing the semantic differences between input pairs of similar but slightly different audio clips. The ADC solves the prob…

Audio captioningDisentanglement

Spoken Dialogue for Information Navigation

2018-07-01 · WS 2018 7 · Alex Papangelis, ros, Panagiotis Papadakos, Yannis Stylianou 외

Aiming to expand the current research paradigm for training conversational AI agents that can address real-world challenges, we take a step away from traditional slot-filling goal-oriented spoken dialogue systems (SDS) a…

Information Retrievalslot-fillingSlot FillingSpoken Dialogue Systems

Multi-Modal Relational Graph for Cross-Modal Video Moment Retrieval

2021-06-19 · CVPR 2021 1 · Yawen Zeng, Da Cao, Xiaochi Wei, Meng Liu 외

Given an untrimmed video and a query sentence, cross-modal video moment retrieval aims to rank a video moment from pre-segmented video moment candidates that best matches the query sentence. Pioneering work typically…

Cross-Modal RetrievalGraph MatchingMoment RetrievalRelation+2