paper-with-me

홈 › Papers

Le MuMo JEPA: Multi-Modal Self-Supervised Representation Learning with Learnable Fusion Tokens

2026-03-25 · Ciem Cornelissen, Sam Leroux, Pieter Simoens arxiv

Self-supervised learning has emerged as a powerful paradigm for learning visual representations without manual annotations, yet most methods still operate on a single modality and therefore miss the complementary structure available from heterogeneous sensors. We present Le MuMo JEPA, a self-supervised framework that learns unified representations from RGB images and aligned companion modalities. In our driving experiments, the second modality is camera-aligned LiDAR depth; we also evaluate RGB-thermal training and transfer on the Teledyne FLIR ADAS benchmark. Our approach extends LeJEPA to the multi-modal setting by learning fusion tokens that act as a latent bottleneck between modality-specific patch stems inside a shared transformer. Our default model employs a pruned fusion strategy: after an initial cross-modal attention layer, modality-specific tokens are dropped, forcing cross-modal information into the shared fusion-token grid as an efficient latent bottleneck before Sketched Isotropic Gaussian Regularization (SIGReg) is applied to the joint multimodal CLS embedding. On Waymo, Le MuMo JEPA gives the strongest performance-efficiency trade-off on downstream patch probes among the from-scratch multimodal baselines, improving CenterNet detection and dense depth while remaining competitive on segmentation. Under from-scratch training on nuScenes, Le MuMo JEPA remains the strongest model, and it also gives the best FLIR results, especially after Waymo-initialized fine-tuning. It also retains the best overall accuracy-efficiency balance in our study at substantially lower compute, memory, and estimated training time.

📄 PDF Abstract BibTeX arXiv:2603.24327

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningRepresentation Learning

Similar Papers 제목 키워드 기반

GeoJEPA: Towards Eliminating Augmentation- and Sampling Bias in Multimodal Geospatial Learning

2025-02-25 · Theodor Lundqvist, Ludvig Delvret

Existing methods for self-supervised representation learning of geospatial regions and map entities rely extensively on the design of pretext tasks, often involving augmentations or heuristic sampling of positive and neg…

Representation Learning

M3-Jepa: Multimodal Alignment via Multi-directional MoE based on the JEPA framework

2024-09-09 · Hongyang Lei, Xiaolong Cheng, Dan Wang, Kun Fan 외

Current multimodal alignment strategies primarily use single or unified modality encoders, while optimizing the alignment on the original token space. Such a framework is easy to implement and incorporate with the pretra…

Computational EfficiencyCross-Modal RetrievalMixture-of-ExpertsQuestion Answering+3

Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning

2025-10-24 · Zihao Jing, Yan Sun, Yan Yi Li, Sugitha Janarthanan 외 arxiv

Multimodal molecular models often suffer from 3D conformer unreliability and modality collapse, limiting their robustness and generalization. We propose MuMo, a structured multimodal fusion framework that addresses these…

Representation Learning

TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems

2025-03-09 · Khang H. N. Vo, Duc P. T. Nguyen, Thong Nguyen, Tho T. Quan

This paper focuses on multimodal alignment within the realm of Artificial Intelligence, particularly in text and image modalities. The semantic gap between the textual and visual modality poses a discrepancy problem towa…

Multimodal Sentiment AnalysisQuestion AnsweringSelf-Supervised LearningSentiment Analysis+1

Self-supervised learning of imaging and clinical signatures using a multimodal joint-embedding predictive architecture

2025-09-18 · Thomas Z. Li, Aravind R. Krishnan, Lianrui Zuo, John M. Still 외 arxiv

The development of multimodal models for pulmonary nodule diagnosis is limited by the scarcity of labeled data and the tendency for these models to overfit on the training distribution. In this work, we leverage self-sup…

Self-Supervised Learning