paper-with-me

홈 › Papers

Dynamic Reflections: Probing Video Representations with Text Alignment

2025-11-04 · Tyler Zhu, Tengda Han, Leonidas Guibas, Viorica Pătrăucean, Maks Ovsjanikov arxiv

The alignment of representations from different modalities has recently been shown to provide insights on the structural similarities and downstream capabilities of different encoders across diverse data types. While significant progress has been made in aligning images with text, the temporal nature of video data remains largely unexplored in this context. In this work, we conduct the first comprehensive study of video-text representation alignment, probing the capabilities of modern video and language encoders. Our findings reveal several key insights. First, we demonstrate that cross-modal alignment highly depends on the richness of both visual (static images vs. multi-frame videos) and text (single caption vs. a collection) data provided at test time, especially when using state-of-the-art video encoders. We propose parametric test-time scaling laws that capture this behavior and show remarkable predictive power against empirical observations. Secondly, we investigate the correlation between semantic alignment and performance on both semantic and non-semantic downstream tasks, providing initial evidence that strong alignment against text encoders may be linked to general-purpose video representation and understanding. Finally, we correlate temporal reasoning with cross-modal alignment providing a challenging test-bed for vision and language models. Overall, our work introduces video-text alignment as an informative zero-shot way to probe the representation power of different encoders for spatio-temporal data. Project page can be found at https://video-prh.github.io/

📄 PDF Abstract BibTeX arXiv:2511.02767

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning

2026-06-19 · Awais Rauf, Ahmed Hasssan, Greg Slabaugh arxiv

Understanding long videos requires fine-grained perception and multi-step, higher-order reasoning over complex, long-range spatio-temporal dynamics. Vision-language models (VLMs) encode video frames into visual tokens an…

Relational ReasoningSemantic Retrieval

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis

2026-06-08 · Samuele Punzo, Niccolò Caselli, Ippokratis Pantelidis, Francesco Massafra 외 arxiv

We study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how this information varies across model families, layers, and probe types. Using frozen-featu…

Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation?

2026-06-02 · Luc P. J. Sträter, Hazel Doughty arxiv

Parameter-efficient fine-tuning (PEFT) and probing enable adaptation of foundation models using only a small number of trainable parameters, making it attractive for video understanding where annotation and computation a…

parameter-efficient fine-tuning

Latent Space Probing for Adult Content Detection in Video Generative Models

2026-04-25 · Alizishaan Khatri, Chiquita Prabhu arxiv

The rapid proliferation of AI-powered video generation systems has introduced significant challenges in content moderation, particularly with respect to adult and sexually explicit material. Existing detection methods op…

Video Generation

What Makes Video World Model Latents Action-Relevant: Prediction over Reconstruction

2026-06-05 · Jewon Yeom, Hanseul Kim, Jeongjae Park, Sungmok Jung 외 arxiv

Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces. We study this question thro…