Semi-supervised multimodal coreference resolution in image narrations
In this paper, we study multimodal coreference resolution, specifically where a longer descriptive text, i.e., a narration is paired with an image. This poses significant challenges due to fine-grained image-text alignment, inherent ambiguity present in narrative language, and unavailability of large annotated training sets. To tackle these challenges, we present a data efficient semi-supervised approach that utilizes image-narration pairs to resolve coreferences and narrative grounding in a multimodal context. Our approach incorporates losses for both labeled and unlabeled data within a cross-modal framework. Our evaluation shows that the proposed approach outperforms strong baselines both quantitatively and qualitatively, for the tasks of coreference resolution and narrative grounding.
Code (1)
Tasks
coreference-resolutionCoreference ResolutionDescriptiveSimilar Papers 제목 키워드 기반
Coreference by Appearance: Visually Grounded Event Coreference Resolution
Event coreference resolution is critical to understand events in the growing number of online news with multiple modalities including text, video, speech, etc. However, the events and entities depicting in different moda…
coreference-resolutionCoreference ResolutionEvent Coreference ResolutionMachine Translation+1Exploring Semi-Supervised Coreference Resolution of Medical Concepts using Semantic and Temporal Features
Multimodal Cross-Document Event Coreference Resolution Using Linear Semantic Transfer and Mixed-Modality Ensembles
Event coreference resolution (ECR) is the task of determining whether distinct mentions of events within a multi-document corpus are actually linked to the same underlying occurrence. Images of the events can help facili…
coreference-resolutionCoreference ResolutionEvent Coreference ResolutionTowards Coreference Resolution for Early Irish
In this article, we present an outline of some of the issues involved in developing a semi-supervised procedure for coreference resolution for early Irish as part of a wider enterprise to create a parsed corpus of histor…
coreference-resolutionCoreference ResolutionGRAVL-BERT: Graphical Visual-Linguistic Representations for Multimodal Coreference Resolution
Learning from multimodal data has become a popular research topic in recent years. Multimodal coreference resolution (MCR) is an important task in this area. MCR involves resolving the references across different modalit…
coreference-resolutionCoreference ResolutionVisual Grounding