paper-with-me

홈 › Papers

Few-Shot Relation Extraction with Hybrid Visual Evidence

2024-03-01 · Jiaying Gong, Hoda Eldardiry

The goal of few-shot relation extraction is to predict relations between name entities in a sentence when only a few labeled instances are available for training. Existing few-shot relation extraction methods focus on uni-modal information such as text only. This reduces performance when there are no clear contexts between the name entities described in text. We propose a multi-modal few-shot relation extraction model (MFS-HVE) that leverages both textual and visual semantic information to learn a multi-modal representation jointly. The MFS-HVE includes semantic feature extractors and multi-modal fusion components. The MFS-HVE semantic feature extractors are developed to extract both textual and visual features. The visual features include global image features and local object features within the image. The MFS-HVE multi-modal fusion unit integrates information from various modalities using image-guided attention, object-guided attention, and hybrid feature attention to fully capture the semantic interaction between visual regions of images and relevant texts. Extensive experiments conducted on two public datasets demonstrate that semantic visual information significantly improves the performance of few-shot relation prediction.

📄 PDF Abstract BibTeX arXiv:2403.00724

Code (0)

등록된 구현이 없습니다.

Tasks

RelationRelation ExtractionRelation PredictionSentence

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Chain of Thought with Explicit Evidence Reasoning for Few-shot Relation Extraction

2023-11-10 · Xilai Ma, Jing Li, Min Zhang

Few-shot relation extraction involves identifying the type of relationship between two specific entities within a text, using a limited number of annotated samples. A variety of solutions to this problem have emerged by …

In-Context LearningMeta-LearningRelationRelation Extraction

DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Models

2026-03-04 · Yangfu Li, Hongjian Zhan, Jiawei Chen, Yuning Gong 외 arxiv

Humans can robustly localize visual evidence and provide grounded answers even in noisy environments by identifying critical cues and then relating them to the full context in a bottom-up manner. Inspired by this, we pro…

Bridging Text and Knowledge with Multi-Prototype Embedding for Few-Shot Relational Triple Extraction

2020-10-30 · COLING 2020 8 · Haiyang Yu, Ningyu Zhang, Shumin Deng, Hongbin Ye 외

Current supervised relational triple extraction approaches require huge amounts of labeled data and thus suffer from poor performance in few-shot settings. However, people can grasp new knowledge by learning a few instan…

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

2026-05-19 · Yujie Wei, Yujin Han, Zhekai Chen, Yongming Li 외 arxiv

Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However, evaluating such frontier models remains a fundamental challenge. Ex…

Video Generation

Structured Semantic Information Helps Retrieve Better Examples for In-Context Learning Applied to Few-Shot Relation Extraction

2026-01-28 · Aunabil Chakma, Mihai Surdeanu, Eduardo Blanco arxiv

This paper presents several strategies to automatically obtain additional examples for in-context learning, effectively transforming relation extraction from a 1-shot to a few-shot setting. Specifically, we introduce a n…

Relation Extraction