paper-with-me

Papers

Object-centric LeJEPA

2026-07-02 · Jakob Geusen, Ender Konukoglu arxiv

Image encoders trained with LeJEPA can deliver strong features for downstream tasks, but, like other image-level self-supervised methods, typically require large training datasets. Aligning representations at the level of objects rather than whole scenes promises greater data efficiency, but doing this in a completely self-supervised way, effectively jointly partitioning a scene and representing its objects, is unstable: the two are locked in a cyclic dependency, partitioning requires meaningful representations, while meaningful representations require consistent partitioning. We sidestep this instability by taking object masks as given during training, using cheap, off-the-shelf SAM proposals. We extend LeJEPA - whose distributional anti-collapse objective ports naturally from whole images to variable-sized sets of objects - to align object-centric representations rather than whole images. An additional instance-separating loss, which treats other objects in the same scene as negatives, further boosts downstream performance. Across two model scales and 10-100% of COCO, object-level LeJEPA outperforms image-level LeJEPA on tracking (DAVIS), classification (ImageNet-1k), segmentation (ADE20k), and re-identification (NAVI).

📄 PDF Abstract BibTeX arXiv:2607.02404

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics

2025-11-11 · Randall Balestriero, Yann LeCun arxiv

Learning manipulable representations of the world and its dynamics is central to AI. Joint-Embedding Predictive Architectures (JEPAs) offer a promising blueprint, but lack of practical guidance and theory has led to ad-h…

Self-Supervised Learning

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

2026-08-27 · Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec 외 arxiv

Video carries the temporal structure of the physical world, yet learning representations from it has remained computationally expensive: prevailing self-supervised methods either prevent representation collapse through a…

Gaussian-Constrained LeJEPA Representations for Unsupervised Scene Discovery and Pose Consistency

2026-01-31 · Mohsen Mostafa arxiv

Unsupervised 3D scene reconstruction from unstructured image collections remains a fundamental challenge in computer vision, particularly when images originate from multiple unrelated scenes and contain significant visua…

Self-Supervised LearningCamera Pose EstimationImage Matching

Laya: A LeJEPA Approach to EEG via Latent Prediction over Reconstruction

2026-03-17 · Saarang Panchavati, Uddhav Panchavati, Hiroki Nariai, Corey Arnold 외 arxiv

Electroencephalography (EEG) is a widely used tool for studying brain function, with applications in clinical neuroscience, diagnosis, and brain-computer interfaces (BCIs). Recent EEG foundation models trained on large u…

Self-Supervised Learning

SPHERE-JEPA: Spherical Prediction with Homogeneous Embeddings

2026-05-26 · Léo Nicollier, Max Dunitz, Marc Pic, Pablo Musé 외 arxiv

A fundamental open question in self-supervised learning (SSL) is the explicit characterization of the optimal geometry of the learned representations. Recently, LeJEPA identified isotropic Gaussian embeddings as optimal …

Self-Supervised Learning