paper-with-me

Papers

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition

2026-04-25 · Balaji Darur, Amanmeet Garg, Makarand Tapaswi arxiv

Video Situation Recognition (VidSitu) addresses the challenging problem of "who did what to whom, with what, how, and where" in a video. It tests thorough video understanding by requiring identification of salient actions and associated short descriptions for event roles across multiple events. Grounding with VidSitu requires spatio-temporal localization of key entities across shots and varied appearances. We posit that coherent video understanding requires consistent identification of entities that play different roles. We propose Multimodal Entity Coreference (MEC) to unite entity descriptions in text with grounding across the video. Towards this, we introduce CineMEC, a multi-stage approach that unites event role mention groups with visual clusters of entities, without explicit grounding supervision during training. Our approach is designed to exploit the synergy between visual grounding and captioning, where improving one influences the other and vice versa. For evaluation, we extend the VidSitu dataset with grounding annotations. While previous work focuses primarily on descriptions, CineMEC improves consistency across both: captioning (+2.5% CIDEr, +7% LEA) and visual grounding (+18% HOTA).

📄 PDF Abstract BibTeX arXiv:2604.23173

Code (0)

등록된 구현이 없습니다.

Tasks

Situation RecognitionVisual Grounding

Similar Papers 제목 키워드 기반

Annotating Near-Identity from Coreference Disagreements

2012-05-01 · LREC 2012 5 · Marta Recasens, M. Ant{\`o}nia Mart{\'\i}, Constantin Orasan

We present an extension of the coreference annotation in the English NP4E and the Catalan AnCora-CA corpora with near-identity relations, which are borderline cases of coreference. The annotated subcorpora have 50K token…

Qualitative and Quantitative Analysis of Diversity in Cross-document Coreference Resolution Datasets

2021-10-16 · ACL ARR October 2021 10 · Anonymous

Established cross-document coreference resolution (CDCR) datasets contain manually annotated event-centric mentions of events and entities that form coreference chains with identity relations. In this paper, we qualitati…

coreference-resolutionCoreference ResolutionCross Document Coreference ResolutionDiversity

Identity and Granularity of Events in Text

2017-04-13 · Piek Vossen, Agata Cybulska

In this paper we describe a method to detect event descrip- tions in different news articles and to model the semantics of events and their components using RDF representations. We compare these descriptions to solve a c…

ArticlesEvent Detection

Cross-document Event Identity via Dense Annotation

2021-09-14 · CoNLL (EMNLP) 2021 11 · Adithya Pratapa, Zhengzhong Liu, Kimihiro Hasegawa, Linwei Li 외

In this paper, we study the identity of textual events from different documents. While the complex nature of event identity is previously studied (Hovy et al., 2013), the case of events across documents is unclear. Prior…

Scoring Coreference Chains with Split-Antecedent Anaphors

2022-05-24 · Silviu Paun, Juntao Yu, Nafise Sadat Moosavi, Massimo Poesio

Anaphoric reference is an aspect of language interpretation covering a variety of types of interpretation beyond the simple case of identity reference to entities introduced via nominal expressions covered by the traditi…