paper-with-me

홈 › Papers

TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems

2025-03-09 · Khang H. N. Vo, Duc P. T. Nguyen, Thong Nguyen, Tho T. Quan

This paper focuses on multimodal alignment within the realm of Artificial Intelligence, particularly in text and image modalities. The semantic gap between the textual and visual modality poses a discrepancy problem towards the effectiveness of multi-modalities fusion. Therefore, we introduce Text-Image Joint Embedding Predictive Architecture (TI-JEPA), an innovative pre-training strategy that leverages energy-based model (EBM) framework to capture complex cross-modal relationships. TI-JEPA combines the flexibility of EBM in self-supervised learning to facilitate the compatibility between textual and visual elements. Through extensive experiments across multiple benchmarks, we demonstrate that TI-JEPA achieves state-of-the-art performance on multimodal sentiment analysis task (and potentially on a wide range of multimodal-based tasks, such as Visual Question Answering), outperforming existing pre-training methodologies. Our findings highlight the potential of using energy-based framework in advancing multimodal fusion and suggest significant improvements for downstream applications.

📄 PDF Abstract BibTeX arXiv:2503.06380

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Sentiment AnalysisQuestion AnsweringSelf-Supervised LearningSentiment AnalysisVisual Question Answering

Methods 이 논문이 사용한 방법론

EBM 설명 없음

Similar Papers 제목 키워드 기반

Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning

2024-10-25 · Shentong Mo, Shengbang Tong

In recent advancements in unsupervised visual representation learning, the Joint-Embedding Predictive Architecture (JEPA) has emerged as a significant method for extracting visual features from unlabeled imagery through …

Representation LearningSelf-Supervised Learning

HEP-JEPA: A foundation model for collider physics using joint embedding predictive architecture

2025-02-06 · Jai Bardhan, Radhikesh Agrawal, Abhiram Tilak, Cyrin Neeraj 외

We present a transformer architecture-based foundation model for tasks at high-energy particle colliders such as the Large Hadron Collider. We train the model to classify jets using a self-supervised strategy inspired by…

Intrinsic-Energy Joint Embedding Predictive Architectures Induce Quasimetric Spaces

2026-02-12 · Anthony Kobanda, Waris Radji arxiv

Joint-Embedding Predictive Architectures (JEPAs) aim to learn representations by predicting target embeddings from context embeddings, inducing a scalar compatibility energy in a latent space. In contrast, Quasimetric Re…

Reinforcement Learning

Denoising with a Joint-Embedding Predictive Architecture

2024-10-02 · Dengsheng Chen, Jie Hu, Xiaoming Wei, Enhua Wu

Joint-embedding predictive architectures (JEPAs) have shown substantial promise in self-supervised representation learning, yet their application in generative modeling remains underexplored. Conversely, diffusion models…

DenoisingImage GenerationRepresentation Learning

Distributed JEPA: A Self-Supervised Framework for Energy Forecasting

2026-09-15 · Liana Toderean, Tudor Cioara, Vasilis Michalakopoulos, Efstathios Sarantinopoulos 외 arxiv

Traditional energy forecasting solutions rely on task-specific supervision and energy asset representations, limiting transferability and the ability to capture general temporal dynamics across heterogeneous assets. We a…

Self-Supervised Learning