paper-with-me

Papers

Progressively Modality Freezing for Multi-Modal Entity Alignment

2024-07-23 · Yani Huang, Xuefeng Zhang, Richong Zhang, Junfan Chen, Jaein Kim

Multi-Modal Entity Alignment aims to discover identical entities across heterogeneous knowledge graphs. While recent studies have delved into fusion paradigms to represent entities holistically, the elimination of features irrelevant to alignment and modal inconsistencies is overlooked, which are caused by inherent differences in multi-modal features. To address these challenges, we propose a novel strategy of progressive modality freezing, called PMF, that focuses on alignmentrelevant features and enhances multi-modal feature fusion. Notably, our approach introduces a pioneering cross-modal association loss to foster modal consistency. Empirical evaluations across nine datasets confirm PMF's superiority, demonstrating stateof-the-art performance and the rationale for freezing modalities. Our code is available at https://github.com/ninibymilk/PMF-MMEA.

📄 PDF Abstract BibTeX arXiv:2407.16168

Code (1)

ninibymilk/pmf-mmea 공식 구현 pytorch

Tasks

Entity AlignmentKnowledge GraphsMulti-modal Entity Alignment

Similar Papers 제목 키워드 기반

CMCC-ReID: Cross-Modality Clothing-Change Person Re-Identification

2026-04-03 · Haoxuan Xu, Hanzi Wang, Guanglin Niu arxiv

Person Re-Identification (ReID) faces severe challenges from modality discrepancy and clothing variation in long-term surveillance scenario. While existing studies have made significant progress in either Visible-Infrare…

Person Re-Identification

Zipper: A Multi-Tower Decoder Architecture for Fusing Modalities

2024-05-29 · Vicky Zayats, Peter Chen, Melissa Ferrari, Dirk Padfield

Integrating multiple generative foundation models, especially those trained on different modalities, into something greater than the sum of its parts poses significant challenges. Two key hurdles are the availability of …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+4

VINO: A Unified Visual Generator with Interleaved OmniModal Context

2026-01-05 · Junyi Chen, Tong He, Zhoujie Fu, Pengfei Wan 외 arxiv

We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independent modules for each modality, VINO uses a…

Instruction FollowingVideo Generation

MEAformer: Multi-modal Entity Alignment Transformer for Meta Modality Hybrid

2022-12-29 · Zhuo Chen, Jiaoyan Chen, Wen Zhang, Lingbing Guo 외

Multi-modal entity alignment (MMEA) aims to discover identical entities across different knowledge graphs (KGs) whose entities are associated with relevant images. However, current MMEA algorithms rely on KG-level modali…

Entity AlignmentKnowledge GraphsMulti-modal Entity Alignment

Relation-Aware Distribution Representation Network for Person Clustering with Multiple Modalities

2023-08-01 · Kaijian Liu, Shixiang Tang, Ziyue Li, Zhishuai Li 외

Person clustering with multi-modal clues, including faces, bodies, and voices, is critical for various tasks, such as movie parsing and identity-based movie editing. Related methods such as multi-view clustering mainly p…

ClusteringRelation