paper-with-me

Papers

Gaussian Joint Embeddings For Self-Supervised Representation Learning

2026-03-26 · Yongchao Huang arxiv

Self-supervised representation learning often relies on deterministic predictive architectures to align context and target views in latent space. While effective in many settings, such methods are limited in genuinely multi-modal inverse problems, where squared-loss prediction collapses towards conditional averages, and they frequently depend on architectural asymmetries to prevent representation collapse. In this work, we propose a probabilistic alternative based on generative joint modeling. We introduce Gaussian Joint Embeddings (GJE) and its multi-modal extension, Gaussian Mixture Joint Embeddings (GMJE), which model the joint density of context and target representations and replace black-box prediction with closed-form conditional inference under an explicit probabilistic model. This yields principled uncertainty estimates and a covariance-aware objective for controlling latent geometry. We further identify a failure mode of naive empirical batch optimization, which we term the Mahalanobis Trace Trap, and develop several remedies spanning parametric, adaptive, and non-parametric settings, including prototype-based GMJE, conditional Mixture Density Networks (GMJE-MDN), topology-adaptive Growing Neural Gas (GMJE-GNG), and a Sequential Monte Carlo (SMC) memory bank. In addition, we show that standard contrastive learning can be interpreted as a degenerate non-parametric limiting case of the GMJE framework. Experiments on synthetic multi-modal alignment tasks and vision benchmarks show that GMJE recovers complex conditional structure, learns competitive discriminative representations, and defines latent densities that are better suited to unconditional sampling than deterministic or unimodal baselines.

📄 PDF Abstract BibTeX arXiv:2603.26799

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningContrastive Learning

Similar Papers 제목 키워드 기반

Gaussian-Constrained LeJEPA Representations for Unsupervised Scene Discovery and Pose Consistency

2026-01-31 · Mohsen Mostafa arxiv

Unsupervised 3D scene reconstruction from unstructured image collections remains a fundamental challenge in computer vision, particularly when images originate from multiple unrelated scenes and contain significant visua…

Self-Supervised LearningCamera Pose EstimationImage Matching

Joint Wasserstein Autoencoders for Aligning Multimodal Embeddings

2019-09-14 · Shweta Mahajan, Teresa Botschen, Iryna Gurevych, Stefan Roth

One of the key challenges in learning joint embeddings of multiple modalities, e.g. of images and text, is to ensure coherent cross-modal semantics that generalize across datasets. We propose to address this through join…

Cross-Modal RetrievalRetrieval

Self-Supervised Slice-to-Volume Reconstruction with Gaussian Representations for Fetal MRI

2026-01-30 · Yinsong Wang, Thomas Fletcher, Xinzhe Luo, Aine Travers Dineen 외 arxiv

Reconstructing 3D fetal MR volumes from motion-corrupted stacks of 2D slices is a crucial and challenging task. Conventional slice-to-volume reconstruction (SVR) methods are time-consuming and require multiple orthogonal…

Why Self-Supervised Encoders Want to Be Normal

2026-04-30 · Yuval Domb arxiv

Self-supervised learning has achieved remarkable empirical success in learning robust representations without explicit labels, most recently demonstrated within the framework of Joint-Embedding Predictive Architectures (…

Self-Supervised Learning

TaxoBell: Gaussian Box Embeddings for Self-Supervised Taxonomy Expansion

2026-01-14 · Sahil Mishra, Srinitish Srinivasan, Srikanta Bedathur, Tanmoy Chakraborty arxiv

Taxonomies form the backbone of structured knowledge representation across diverse domains, enabling applications such as e-commerce and semantic search. Yet, manual taxonomy expansion is labor-intensive and slow. Existi…