Visual Attention Model for Name Tagging in Multimodal Social Media
Everyday billions of multimodal posts containing both images and text are shared in social media sites such as Snapchat, Twitter or Instagram. This combination of image and text in a single message allows for more creative and expressive forms of communication, and has become increasingly common in such sites. This new paradigm brings new challenges for natural language understanding, as the textual component tends to be shorter, more informal, and often is only understood if combined with the visual context. In this paper, we explore the task of name tagging in multimodal social media posts. We start by creating two new multimodal datasets: the first based on Twitter posts and the second based on Snapchat captions (exclusively submitted to public and crowd-sourced stories). We then propose a novel model architecture based on Visual Attention that not only provides deeper visual understanding on the decisions of the model, but also significantly outperforms other state-of-the-art baseline methods for this task.
Code (0)
등록된 구현이 없습니다.
Tasks
Natural Language UnderstandingQuestion AnsweringSimilar Papers 제목 키워드 기반
A Multimodal Framework for Video Ads Understanding
There is a growing trend in placing video advertisements on social platforms for online marketing, which demands automatic approaches to understand the contents of advertisements effectively. Taking the 2021 TAAC competi…
MarketingOptical Character RecognitionOptical Character Recognition (OCR)Scene Segmentation+3TriMod Fusion for Multimodal Named Entity Recognition in Social Media
Social media platforms serve as invaluable sources of user-generated content, offering insights into various aspects of human behavior. Named Entity Recognition (NER) plays a crucial role in analyzing such content by ide…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERMultimodal Named Entity Recognition for Short Social Media Posts
We introduce a new task called Multimodal Named Entity Recognition (MNER) for noisy user-generated data such as tweets or Snapchat captions, which comprise short text with accompanying images. These social media posts of…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERJoint Multimodal Entity-Relation Extraction Based on Edge-enhanced Graph Alignment Network and Word-pair Relation Tagging
Multimodal named entity recognition (MNER) and multimodal relation extraction (MRE) are two fundamental subtasks in the multimodal knowledge graph construction task. However, the existing methods usually handle two tasks…
graph constructionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+3Aiding Intra-Text Representations with Visual Context for Multimodal Named Entity Recognition
With massive explosion of social media such as Twitter and Instagram, people daily share billions of multimedia posts, containing images and text. Typically, text in these posts is short, informal and noisy, leading to a…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)