paper-with-me

Papers

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis

2026-06-08 · Samuele Punzo, Niccolò Caselli, Ippokratis Pantelidis, Francesco Massafra, Salvatore Lo Sardo, Mohammadreza Salehi arxiv

We study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how this information varies across model families, layers, and probe types. Using frozen-feature probing on IntPhys2 and Minimal Video Pairs (MVP), we compare predictive joint-embedding models (V-JEPA), masked reconstruction models (VideoMAE), and a diffusion-based video generator (LTX-Video). V-JEPA achieves the strongest overall results across benchmarks, especially with probes that model temporal dynamics, while VideoMAE remains competitive and LTX-Video recovers weaker but non-trivial signal. Layerwise analyses show that physics-relevant information is weakest in early layers and becomes most accessible at intermediate-to-late depth, and temporal controls show that disrupting frame order substantially reduces performance, especially on MVP. Together, these results suggest that intuitive-physics knowledge emerges reliably in pretrained video representations, but its accessibility depends strongly on pretraining paradigm, representational depth, and readout mechanism.

📄 PDF Abstract BibTeX arXiv:2606.09646

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

2025-05-29 · Xiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng 외

Recent advancements in text-to-video (T2V) diffusion models have enabled high-fidelity and realistic video synthesis. However, current T2V models often struggle to generate physically plausible content due to their limit…

Self-Supervised LearningVideo GenerationVideo Understanding

Intuitive physics understanding emerges from self-supervised pretraining on natural videos

2025-02-17 · Quentin Garrido, Nicolas Ballas, Mahmoud Assran, Adrien Bardes 외

We investigate the emergence of intuitive physics understanding in general-purpose deep neural network models trained to predict masked regions in natural videos. Leveraging the violation-of-expectation framework, we fin…

Video Prediction

LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference

2025-10-13 · Jianhao Yuan, Fabio Pizzati, Francesco Pinto, Lars Kunze 외 arxiv

Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due …

Interpreting Physics in Video World Models

2026-02-04 · Sonia Joseph, Quentin Garrido, Randall Balestriero, Matthew Kowal 외 arxiv

A long-standing question in physical reasoning is whether video-based models need to rely on factorized representations of physical variables in order to make physically accurate predictions, or whether they can implicit…

Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models

2026-03-30 · Nanxi Li, Xiang Wang, Yuanjie Chen, Haode Zhang 외 arxiv

While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focu…