paper-with-me

Papers

Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video

2026-03-14 · Yuting Tan, Xilong Cheng, Yunxiao Qin, Zhengnan Li, Jingjing Zhang arxiv

Humans develop visual intelligence through perceiving and interacting with their environment - a self-supervised learning process grounded in egocentric experience. Inspired by this, we ask how can artificial systems learn stable object representations from continuous, uncurated first-person videos without relying on manual annotations. This setting poses challenges of separating, recognizing, and persistently tracking objects amid clutter, occlusion, and ego-motion. We propose EgoViT, a unified vision Transformer framework designed to learn stable object representations from unlabeled egocentric video. EgoViT bootstraps this learning process by jointly discovering and stabilizing "proto-objects" through three synergistic mechanisms: (1) Proto-object Learning, which uses intra-frame distillation to form discriminative representations; (2) Depth Regularization, which grounds these representations in geometric structure; and (3) Teacher-Filtered Temporal Consistency, which enforces identity over time. This creates a virtuous cycle where initial object hypotheses are progressively refined into stable, persistent representations. The framework is trained end-to-end on unlabeled first-person videos and exhibits robustness to geometric priors of varied origin and quality. On standard benchmarks, EgoViT achieves +8.0% CorLoc improvement in unsupervised object discovery and +4.8% mIoU improvement in semantic segmentation, demonstrating its potential to lay a foundation for robust visual abstraction in embodied intelligence.

📄 PDF Abstract BibTeX arXiv:2603.13912

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningSemantic Segmentation

Similar Papers 제목 키워드 기반

Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities

2023-06-07 · NeurIPS 2023 11 · Andrii Zadaianchuk, Maximilian Seitzer, Georg Martius

Unsupervised video-based object-centric learning is a promising avenue to learn structured representations from large, unlabeled video collections, but previous approaches have only managed to scale to real-world dataset…

ObjectObject Discovery

Scaling Up Semi-supervised Learning with Unconstrained Unlabelled Data

2023-06-02 · Shuvendu Roy, Ali Etemad

We propose UnMixMatch, a semi-supervised learning framework which can learn effective representations from unconstrained unlabelled data in order to scale up performance. Most existing semi-supervised methods rely on the…

Image ClassificationNetwork PruningSemi-Supervised Image Classification

STERLING: Self-Supervised Terrain Representation Learning from Unconstrained Robot Experience

2023-09-26 · Haresh Karnan, Elvin Yang, Daniel Farkash, Garrett Warnell 외

Terrain awareness, i.e., the ability to identify and distinguish different types of terrain, is a critical ability that robots must have to succeed at autonomous off-road navigation. Current approaches that provide robot…

Representation LearningVisual Navigation

`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation

2026-07-08 · Waqas Arshid, Mohammad Awrangjeb, Alan Wee-Chung Liew, Yongsheng Gao arxiv

Video object segmentation (VOS) is a fundamental task in video understanding, requiring accurate delineation and consistent tracking of objects across frames. While supervised methods achieve strong performance, they rel…

Video Object SegmentationSelf-Supervised Learning

SimNP: Learning Self-Similarity Priors Between Neural Points

2023-09-07 · ICCV 2023 1 · Christopher Wewer, Eddy Ilg, Bernt Schiele, Jan Eric Lenssen

Existing neural field representations for 3D object reconstruction either (1) utilize object-level representations, but suffer from low-quality details due to conditioning on a global latent code, or (2) are able to perf…

3D Object ReconstructionObjectObject Reconstruction