paper-with-me

Papers

Temporal Slowness in Central Vision Drives Semantic Object Learning

2026-02-04 · Timothy Schaumlöffel, Arthur Aubret, Gemma Roig, Jochen Triesch arxiv

Humans acquire semantic object representations from egocentric visual streams with minimal supervision, but the underlying mechanisms remain unclear. Importantly, the visual system only processes the center of its field of view with high resolution and it learns similar representations for visual inputs occurring close in time. This emphasizes slowly changing information around gaze locations. This study investigates the role of central vision and slowness learning in the formation of semantic object representations from human-like visual experience. We simulate five months of human-like visual experience using the Ego4D dataset and a state-of-the-art gaze prediction model. We extract image crops around predicted gaze locations to train a time-contrastive Self-Supervised Learning model. Our results show that exploiting temporal slowness when learning from central visual field experience improves the encoding of different facets of object semantics. Specifically, focusing on central vision strengthens the extraction of foreground object features, while considering temporal slowness, especially in conjunction with eye movements, allows the model to encode broader semantic information about objects. These findings provide new insights into the mechanisms by which humans may develop semantic object representations from natural visual experience. Our code will be made public upon acceptance. Code is available at https://github.com/t9s9/central-vision-ssl.

📄 PDF Abstract BibTeX arXiv:2602.04462

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

Video-ReTime: Learning Temporally Varying Speediness for Time Remapping

2022-05-11 · Simon Jenni, Markus Woodson, Fabian Caba Heilbron

We propose a method for generating a temporally remapped video that matches the desired target duration while maximally preserving natural video dynamics. Our approach trains a neural network through self-supervision to …

Action Recognition

Unsupervised Feature Learning from Temporal Data

2015-04-09 · Ross Goroshin, Joan Bruna, Jonathan Tompson, David Eigen 외

Current state-of-the-art classification and detection algorithms rely on supervised training. In this work we study unsupervised feature learning in the context of temporally coherent video data. We focus on feature lear…

General ClassificationMetric Learning

Unsupervised Learning of Spatiotemporally Coherent Metrics

2014-12-18 · ICCV 2015 12 · Ross Goroshin, Joan Bruna, Jonathan Tompson, David Eigen 외

Current state-of-the-art classification and detection algorithms rely on supervised training. In this work we study unsupervised feature learning in the context of temporally coherent video data. We focus on feature lear…

General ClassificationMetric Learning

Autoencoding Slow Representations for Semi-supervised Data Efficient Regression

2020-12-11 · Oliver Struckmeier, Kshitij Tiwari, Ville Kyrki

The slowness principle is a concept inspired by the visual cortex of the brain. It postulates that the underlying generative factors of a quickly varying sensory signal change on a slower time scale. Unsupervised learnin…

regressionRepresentation Learning

Self-taught learning of a deep invariant representation for visual tracking via temporal slowness principle

2016-04-14 · Jason Kuen, Kian Ming Lim, Chin Poo Lee

Visual representation is crucial for a visual tracking method's performances. Conventionally, visual representations adopted in visual tracking rely on hand-crafted computer vision descriptors. These descriptors were dev…

Visual Tracking