paper-with-me

Papers

Multi-level Cross-modal Alignment for Image Clustering

2024-01-22 · Liping Qiu, Qin Zhang, Xiaojun Chen, Shaotian Cai

Recently, the cross-modal pretraining model has been employed to produce meaningful pseudo-labels to supervise the training of an image clustering model. However, numerous erroneous alignments in a cross-modal pre-training model could produce poor-quality pseudo-labels and degrade clustering performance. To solve the aforementioned issue, we propose a novel \textbf{Multi-level Cross-modal Alignment} method to improve the alignments in a cross-modal pretraining model for downstream tasks, by building a smaller but better semantic space and aligning the images and texts in three levels, i.e., instance-level, prototype-level, and semantic-level. Theoretical results show that our proposed method converges, and suggests effective means to reduce the expected clustering risk of our method. Experimental results on five benchmark datasets clearly show the superiority of our new method.

📄 PDF Abstract BibTeX arXiv:2401.11740

Code (0)

등록된 구현이 없습니다.

Tasks

Clusteringcross-modal alignmentImage Clustering

Similar Papers 제목 키워드 기반

EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment

2024-10-08 · Yifei Xing, Xiangyuan Lan, Ruiping Wang, Dongmei Jiang 외

Mamba-based architectures have shown to be a promising new direction for deep learning models owing to their competitive performance and sub-quadratic deployment speed. However, current Mamba multi-modal large language m…

cross-modal alignmentHallucinationMamba

A Three-Level Alignment Framework for Large-Scale 3D Retrieval and Controlled 4D Generation

2025-12-25 · Philip Xu arxiv

We introduce Uni4D, a unified framework for large scale open vocabulary 3D retrieval and controlled 4D generation based on structured three level alignment across text, 3D models, and image modalities. Built upon the Ali…

Text to 3D

MVPTR: Multi-Level Semantic Alignment for Vision-Language Pre-Training via Multi-Stage Learning

2022-01-29 · Zejun Li, Zhihao Fan, Huaixiao Tou, Jingjing Chen 외

Previous vision-language pre-training models mainly construct multi-modal inputs with tokens and objects (pixels) followed by performing cross-modality interaction between them. We argue that the input of only tokens and…

Image-text matchingLanguage ModelingLanguage ModellingMasked Language Modeling+2

Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning

2022-10-12 · Fuying Wang, Yuyin Zhou, Shujun Wang, Varut Vardhanabhuti 외

Learning medical visual representations directly from paired radiology reports has become an emerging topic in representation learning. However, existing medical image-text joint learning methods are limited by instance …

Contrastive Learningcross-modal alignmentimage-classificationImage Classification+4

Cross-Modality Paired-Images Generation for RGB-Infrared Person Re-Identification

2020-02-10 · Guan-An Wang, Tianzhu Zhang. Yang Yang, Jian Cheng, Jianlong Chang 외

RGB-Infrared (IR) person re-identification is very challenging due to the large cross-modality variations between RGB and IR images. The key solution is to learn aligned features to the bridge RGB and IR modalities. Howe…

Person Re-Identification