paper-with-me

Papers

UNITER-Based Situated Coreference Resolution with Rich Multimodal Input

2021-12-07 · Yichen Huang, Yuchen Wang, Yik-Cheung Tam

We present our work on the multimodal coreference resolution task of the Situated and Interactive Multimodal Conversation 2.0 (SIMMC 2.0) dataset as a part of the tenth Dialog System Technology Challenge (DSTC10). We propose a UNITER-based model utilizing rich multimodal context such as textual dialog history, object knowledge base and visual dialog scenes to determine whether each object in the current scene is mentioned in the current dialog turn. Results show that the proposed approach outperforms the official DSTC10 baseline substantially, with the object F1 score boosted from 36.6% to 77.3% on the development set, demonstrating the effectiveness of the proposed object representations from rich multimodal input. Our model ranks second in the official evaluation on the object coreference resolution task with an F1 score of 73.3% after model ensembling.

📄 PDF Abstract BibTeX arXiv:2112.03521

Code (1)

i-need-sleep/mmcoref_cleaned 공식 구현 pytorch

Tasks

coreference-resolutionCoreference ResolutionObjectVisual Dialog

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Annotating anaphoric phenomena in situated dialogue

2021-06-01 · ACL (mmsr, IWCS) 2021 6 · Sharid Loáiciga, Simon Dobnik, David Schlangen

In recent years several corpora have been developed for vision and language tasks. With this paper, we intend to start a discussion on the annotation of referential phenomena in situated dialogue. We argue that there is …

coreference-resolutionCoreference Resolution

Signed Coreference Resolution

2021-11-01 · EMNLP 2021 11 · Kayo Yin, Kenneth DeHaan, Malihe Alikhani

Coreference resolution is key to many natural language processing tasks and yet has been relatively unexplored in Sign Language Processing. In signed languages, space is primarily used to establish reference. Solving cor…

coreference-resolutionCoreference Resolution

Situated and Interactive Multimodal Conversations

2020-06-02 · COLING 2020 8 · Seungwhan Moon, Satwik Kottur, Paul A. Crook, Ankita De 외

Next generation virtual assistants are envisioned to handle multimodal inputs (e.g., vision, memories of previous interactions, in addition to the user's utterances), and perform multimodal actions (e.g., displaying a ro…

Response Generation

Exploring Multi-Modal Representations for Ambiguity Detection & Coreference Resolution in the SIMMC 2.0 Challenge

2022-02-25 · Javier Chiyah-Garcia, Alessandro Suglia, José Lopes, Arash Eshghi 외

Anaphoric expressions, such as pronouns and referential descriptions, are situated with respect to the linguistic context of prior turns, as well as, the immediate visual environment. However, a speaker's referential des…

coreference-resolutionCoreference Resolution

Semi-supervised multimodal coreference resolution in image narrations

2023-10-20 · Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen

In this paper, we study multimodal coreference resolution, specifically where a longer descriptive text, i.e., a narration is paired with an image. This poses significant challenges due to fine-grained image-text alignme…

coreference-resolutionCoreference ResolutionDescriptive