paper-with-me

홈 › Papers

Can images help recognize entities? A study of the role of images for Multimodal NER

2020-10-23 · WNUT (ACL) 2021 11 · Shuguang Chen, Gustavo Aguilar, Leonardo Neves, Thamar Solorio

Multimodal named entity recognition (MNER) requires to bridge the gap between language understanding and visual context. While many multimodal neural techniques have been proposed to incorporate images into the MNER task, the model's ability to leverage multimodal interactions remains poorly understood. In this work, we conduct in-depth analyses of existing multimodal fusion techniques from different perspectives and describe the scenarios where adding information from the image does not always boost performance. We also study the use of captions as a way to enrich the context for MNER. Experiments on three datasets from popular social platforms expose the bottleneck of existing multimodal models and the situations where using captions is beneficial.

📄 PDF Abstract BibTeX arXiv:2010.12712

Code (1)

RiTUAL-UH/multimodal_NER 공식 구현 pytorch

Tasks

Image Captioningnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Similar Papers 제목 키워드 기반

CEAR: Automatic construction of a knowledge graph of chemical entities and roles from scientific literature

2024-07-31

Ontologies are formal representations of knowledge in specific domains that provide a structured framework for organizing and understanding complex information. Creating ontologies, however, is a complex and time-consumi…

TextCaps: a Dataset for Image Captioning with Reading Comprehension

2020-03-24 · ECCV 2020 8 · Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, Amanpreet Singh

Image descriptions can help visually impaired people to quickly understand the image content. While we made significant progress in automatically describing images and optical character recognition, current approaches ar…

Image CaptioningOptical Character RecognitionOptical Character Recognition (OCR)Reading Comprehension+1

Learning semantic Image attributes using Image recognition and knowledge graph embeddings

2020-09-12 · Ashutosh Tiwari, Sandeep Varma

Extracting structured knowledge from texts has traditionally been used for knowledge base generation. However, other sources of information, such as images can be leveraged into this process to build more complete and ri…

Graph EmbeddingKnowledge Graph EmbeddingKnowledge Graph EmbeddingsKnowledge Graphs

DAMO-NLP at NLPCC-2022 Task 2: Knowledge Enhanced Robust NER for Speech Entity Linking

2022-09-27 · Shen Huang, Yuchen Zhai, Xinwei Long, Yong Jiang 외

Speech Entity Linking aims to recognize and disambiguate named entities in spoken languages. Conventional methods suffer gravely from the unfettered speech styles and the noisy transcripts generated by ASR systems. In th…

Entity Linkingnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+4

Cross-modal codification of images with auditory stimuli: a language for the visually impaired

2017-05-15

In this study we describe a methodology to realize visual images cognition in the broader sense, by a cross-modal stimulation through the auditory channel. An original algorithm of conversion from bi-dimensional images t…