paper-with-me

홈 › Papers

TAMM: TriAdapter Multi-Modal Learning for 3D Shape Understanding

2024-02-28 · CVPR 2024 1 · Zhihao Zhang, Shengcao Cao, Yu-Xiong Wang

The limited scale of current 3D shape datasets hinders the advancements in 3D shape understanding, and motivates multi-modal learning approaches which transfer learned knowledge from data-abundant 2D image and language modalities to 3D shapes. However, even though the image and language representations have been aligned by cross-modal models like CLIP, we find that the image modality fails to contribute as much as the language in existing multi-modal 3D representation learning methods. This is attributed to the domain shift in the 2D images and the distinct focus of each modality. To more effectively leverage both modalities in the pre-training, we introduce TriAdapter Multi-Modal Learning (TAMM) -- a novel two-stage learning approach based on three synergistic adapters. First, our CLIP Image Adapter mitigates the domain gap between 3D-rendered images and natural images, by adapting the visual representations of CLIP for synthetic image-text pairs. Subsequently, our Dual Adapters decouple the 3D shape representation space into two complementary sub-spaces: one focusing on visual attributes and the other for semantic understanding, which ensure a more comprehensive and effective multi-modal pre-training. Extensive experiments demonstrate that TAMM consistently enhances 3D representations for a wide range of 3D encoder architectures, pre-training datasets, and downstream tasks. Notably, we boost the zero-shot classification accuracy on Objaverse-LVIS from 46.8\% to 50.7\%, and improve the 5-way 10-shot linear probing classification accuracy on ModelNet40 from 96.1\% to 99.0\%. Project page: https://alanzhangcs.github.io/tamm-page.

📄 PDF Abstract BibTeX arXiv:2402.18490

Code (0)

등록된 구현이 없습니다.

Tasks

3D Shape RepresentationRepresentation LearningZero-shot 3D classificationZero-shot 3D Point Cloud Classificationzero-shot-classificationZero-Shot Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Adapter 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation

2025-01-13 · Han Liu, Yinwei Wei, Fan Liu, Wenjie Wang 외

Multimodal information (e.g., visual, acoustic, and textual) has been widely used to enhance representation learning for micro-video recommendation. For integrating multimodal information into a joint representation of m…

Meta-LearningMultimodal RecommendationRepresentation Learning

Text-centric Alignment for Multi-Modality Learning

2024-02-12 · Yun-Da Tsai, Ting-Yu Yen, Pei-Fu Guo, Zhe-Yan Li 외

This research paper addresses the challenge of modality mismatch in multimodal learning, where the modalities available during inference differ from those available at training. We propose the Text-centric Alignment for …

In-Context Learning

Visual Question Answering on Multiple Remote Sensing Image Modalities

2025-05-21 · Hichem Boussaid, Lucrezia Tosato, Flora Weissgerber, Camille Kurtz 외

The extraction of visual features is an essential step in Visual Question Answering (VQA). Building a good visual representation of the analyzed scene is indeed one of the essential keys for the system to be able to corr…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Separating the effects of experimental noise from inherent system variability in voltammetry: the $[$Fe(CN)$_6]^{3-/ 4-}$ process

2018-09-18

Recently, we have introduced the use of techniques drawn from Bayesian statistics to recover kinetic and thermodynamic parameters from voltammetric data, and were able to show that the technique of large amplitude ac vol…

Impact des modalités induites par les outils d’annotation manuelle : exemple de la détection des erreurs de français (Impact of modalities induced by manual annotation tools : example of French error detection)

2022-06-01 · JEP/TALN/RECITAL 2022 6 · Anaëlle Baledent

Certains choix effectués lors de la construction d’une campagne d’annotation peuvent avoir des conséquences sur les annotations produites. En menant une campagne sur la détection des erreurs de français, aux paramètres m…