paper-with-me

Papers

Cross-media Structured Common Space for Multimedia Event Extraction

2020-05-05 · ACL 2020 6 · Manling Li, Alireza Zareian, Qi Zeng, Spencer Whitehead, Di Lu, Heng Ji, Shih-Fu Chang

We introduce a new task, MultiMedia Event Extraction (M2E2), which aims to extract events and their arguments from multimedia documents. We develop the first benchmark and collect a dataset of 245 multimedia news articles with extensively annotated events and arguments. We propose a novel method, Weakly Aligned Structured Embedding (WASE), that encodes structured representations of semantic information from textual and visual data into a common embedding space. The structures are aligned across modalities by employing a weakly supervised training strategy, which enables exploiting available resources without explicit cross-media annotation. Compared to uni-modal state-of-the-art methods, our approach achieves 4.0% and 9.8% absolute F-score gains on text event argument role labeling and visual event extraction. Compared to state-of-the-art multimedia unstructured representations, we achieve 8.3% and 5.0% absolute F-score gains on multimedia event extraction and argument role labeling, respectively. By utilizing images, we extract 21.4% more event mentions than traditional text-only methods.

📄 PDF Abstract BibTeX arXiv:2005.02472

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesEvent Extraction

Similar Papers 제목 키워드 기반

Understanding Social Media Cross-Modality Discourse in Linguistic Space

2023-02-26 · Chunpu Xu, Hanzhuo Tan, Jing Li, Piji Li

The multimedia communications with texts and images are popular on social media. However, limited studies concern how images are structured with texts to form coherent meanings in human cognition. To fill in the gap, we …

16k

Cross-Lingual Cross-Platform Rumor Verification Pivoting on Multimedia Content

2018-08-14 · EMNLP 2018 10 · Weiming Wen, Songwen Su, Zhou Yu

With the increasing popularity of smart devices, rumors with multimedia content become more and more common on social networks. The multimedia information usually makes rumors look more convincing. Therefore, finding an …

Semantic SimilaritySemantic Textual Similarity

MMTB: Evaluating Terminal Agents on Multimedia-File Tasks

2026-05-08 · Chiyeong Heo, Jaechang Kim, Junhyuk Kwon, Hoyoung Kim 외 arxiv

Terminals provide a powerful interface for AI agents by exposing diverse tools for automating complex workflows, yet existing terminal-agent benchmarks largely focus on tasks grounded in text, code, and structured files.…

Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures

2025-10-23 · Rahul Raja, Arpita Vats arxiv

Question Answering (QA) systems have traditionally relied on structured text data, but the rapid growth of multimedia content (images, audio, video, and structured metadata) has introduced new challenges and opportunitie…

Question AnsweringAnswer Generation

Discourse in Multimedia: A Case Study in Information Extraction

2018-11-13 · Mrinmaya Sachan, Kumar Avinava Dubey, Eduard H. Hovy, Tom M. Mitchell 외

To ensure readability, text is often written and presented with due formatting. These text formatting devices help the writer to effectively convey the narrative. At the same time, these help the readers pick up the stru…