paper-with-me

Papers

Modeling Dense Cross-Modal Interactions for Joint Entity-Relation Extraction

2020-07-01 · Shan Zhao, Minghao Hu, Zhiping Cai, Fang Liu

Joint extraction of entities and their relations benefits from the close interaction between named entities and their relation information. Therefore, how to effectively model such cross-modal interactions is critical for the final performance. Previous works have used simple methods such as label-feature concatenation to perform coarse-grained semantic fusion among cross-modal instances, but fail to capture fine-grained correlations over token and label spaces, resulting in insufficient interactions. In this paper, we propose a deep Cross-Modal Attention Network (CMAN) for joint entity and relation extraction. The network is carefully constructed by stacking multiple attention units in depth to fully model dense interactions over token-label spaces, in which two basic attention units are proposed to explicitly capture fine-grained correlations across different modalities (e.g., token-to-token and labelto-token). Experiment results on CoNLL04 dataset show that our model obtains state-of-the-art results by achieving 90.62% F1 on entity recognition and 72.97% F1 on relation classification. In ADE dataset, our model surpasses existing approaches by more than 1.9% F1 on relation classification. Extensive analyses further confirm the effectiveness of our approach.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Joint Entity and Relation ExtractionRelationRelation ClassificationRelation Extraction

Similar Papers 제목 키워드 기반

Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution

2026-05-12 · Chen Wu, Ling Wang, Zhuoran Zheng, Xiangyu Chen 외 arxiv

Guided depth super-resolution (GDSR) reconstructs HR depth maps from LR inputs with HR RGB guidance. Existing methods either model each modality independently or rely on computationally expensive attention mechanisms wit…

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention

2026-06-04 · Giordano Cicchetti, Eleonora Grassucci, Danilo Comminiello arxiv

Transformer-based multimodal models rely on attention mechanisms to integrate information across heterogeneous modalities. Despite their success, existing multimodal attention formulations compute their scores through co…

JointNet: Extending Text-to-Image Diffusion for Dense Distribution Modeling

2023-10-10 · Jingyang Zhang, Shiwei Li, Yuanxun Lu, Tian Fang 외

We introduce JointNet, a novel neural network architecture for modeling the joint distribution of images and an additional dense modality (e.g., depth maps). JointNet is extended from a pre-trained text-to-image diffusio…

Depth EstimationDepth PredictionImage Generation

Euro-PVI: Pedestrian Vehicle Interactions in Dense Urban Centers

2021-06-22 · CVPR 2021 1 · Apratim Bhattacharyya, Daniel Olmeda Reino, Mario Fritz, Bernt Schiele

Accurate prediction of pedestrian and bicyclist paths is integral to the development of reliable autonomous vehicles in dense urban environments. The interactions between vehicle and pedestrian or bicyclist have a signif…

Autonomous VehiclesPedestrian Trajectory PredictionTrajectory Prediction

Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations

2024-03-04 · CVPR 2024 1 · Sangmin Lee, Bolin Lai, Fiona Ryan, Bikram Boote 외

Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However, most prior works on multimodal social cues focus predominantly on single-p…

coreference-resolutionCoreference Resolution