paper-with-me

홈 › Papers

A Matter of Time: Revealing the Structure of Time in Vision-Language Models

2025-10-22 · Nidham Tekaya, Manuela Waldner, Matthias Zeppelzauer arxiv

Large-scale vision-language models (VLMs) such as CLIP have gained popularity for their generalizable and expressive multimodal representations. By leveraging large-scale training data with diverse textual metadata, VLMs acquire open-vocabulary capabilities, solving tasks beyond their training scope. This paper investigates the temporal awareness of VLMs, assessing their ability to position visual content in time. We introduce TIME10k, a benchmark dataset of over 10,000 images with temporal ground truth, and evaluate the time-awareness of 37 VLMs by a novel methodology. Our investigation reveals that temporal information is structured along a low-dimensional, non-linear manifold in the VLM embedding space. Based on this insight, we propose methods to derive an explicit ``timeline'' representation from the embedding space. These representations model time and its chronological progression and thereby facilitate temporal reasoning tasks. Our timeline approaches achieve competitive to superior accuracy compared to a prompt-based baseline while being computationally efficient. All code and data are available at https://tekayanidham.github.io/timeline-page/.

📄 PDF Abstract BibTeX arXiv:2510.19559

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Which Histories Matter for Time Series Forecasting? Learning Predictive Relevance with Future Supervision

2026-08-24 · Yong-Hoon Choi, Youngjin Cho arxiv

Historical retrieval for time-series prediction commonly treats past similarity as a proxy for usefulness. We ask a different question: which historical examples should be expected to matter for a query? We define predic…

Time Series Forecasting

When Does Language Matter? Multilingual Instructions Reveal Step-wise Language Sensitivity in Vision-Language-Action Models

2026-06-10 · Xuan Dong, Zhe Han, Tianhao Niu, Qingfu Zhu 외 arxiv

Vision-Language-Action (VLA) models have shown strong performance in language-conditioned robotic manipulation, yet their robustness to linguistic variation remains poorly understood. In this work, we present the first s…

Masked Autoencoder Pretraining on Strong-Lensing Images for Joint Dark-Matter Model Classification and Super-Resolution

2025-12-07 · Achmad Ardani Prasha, Clavino Ourizqi Rachmadi, Muhamad Fauzan Ibnu Syahlan, Naufal Rahfi Anugerah 외 arxiv

Strong gravitational lensing can reveal the influence of dark-matter substructure in galaxies, but analyzing these effects from noisy, low-resolution images poses a significant challenge. In this work, we propose a maske…

Temporal sequences of brain activity at rest are constrained by white matter structure and modulated by cognitive demands

2019-09-30

A diverse white matter network and finely tuned neuronal membrane properties allow the brain to transition seamlessly between cognitive states. However, it remains unclear how static structural connections guide the temp…

Temporal Sequences

Revealing the 3D Cosmic Web through Gravitationally Constrained Neural Fields

2025-04-21 · Brandon Zhao, Aviad Levis, Liam Connor, Pratul P. Srinivasan 외

Weak gravitational lensing is the slight distortion of galaxy shapes caused primarily by the gravitational effects of dark matter in the universe. In our work, we seek to invert the weak lensing signal from 2D telescope …

3D Reconstruction