paper-with-me

홈 › Papers

Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception

2025-10-01 · Kathy Garcia, Leyla Isik arxiv

Current video foundation models, including the strongest self-supervised models such as V-JEPA2, fail to capture how humans organize social information in dynamic scenes. For example, across a range of diverse vision models tested, none were able to predict human similarity judgments to social video clips as well as a sentence embedding model of the caption text (MPNet). We show this gap in vision model performance can be closed by a compact behavioral supervisory signal. We introduce behavioral geometric supervision (BGS): a hybrid objective that constrains local and global pairwise embedding geometry to match the relational similarity structure across videos. We apply this method using a new human similarity dataset, containing 49,484 odd-one-out judgments from 250 naturalistic social video clips, and low-rank adaptation across four ViT backbones (V-JEPA 2/2.1, TimeSformer, VideoMAE, and CLIP). We find that one of the best fine-tuned models, V-JEPA 2.1, nearly triples in performance compared to the pre-trained baseline and reaches close to the noise ceiling, exceeding the strongest sentence-embedding baseline. In addition, finetuned models (i) capture unique variance in human judgments that caption-based language embeddings do not, (ii) develop interpretable social-affective attributes (valence, arousal, and dominance) despite never being trained on any of these attributes, (iii) zero-shot transfer to a separate dataset of out-of-distribution abstract social interactions, and (iv) shift spatial attention from scene context to socially informative regions (faces, gaze, and interacting bodies). A matched language-distillation control fails to reproduce these gains, ruling out caption transfer as the mechanism. Our results show how a modest amount of human behavioral data can steer video models toward human-like social visual understanding.

📄 PDF Abstract BibTeX arXiv:2510.01502

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video

2026-02-08 · Zihui Gao, Ke Liu, Donny Y. Chen, Duochao Shi 외 arxiv

Geometric foundation models show promise in 3D reconstruction, yet their progress is severely constrained by the scarcity of diverse, large-scale 3D annotations. While Internet videos offer virtually unlimited raw data, …

Zero-shot Generalization3D ReconstructionPoint Clouds

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation

2026-04-27 · Weijie Wang, Xiaoxuan He, Youping Gu, Yifan Yang 외 arxiv

Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often …

Text-to-Video GenerationReinforcement Learning

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation

2026-05-20 · Kaichen Zhou, Yuzhen Chen, Fangneng Zhan, Hang Hua 외 arxiv

Video world models can generate realistic futures from a single instruction, but they often fail to track the same physical points consistently across time. As a result, the generated videos appear plausible, yet lack th…

Robot ManipulationVideo Prediction

PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

2026-07-06 · Sensen Gao, Zhaoqing Wang, Qihang Cao, Dongdong Yu 외 arxiv

3D reconstruction and generation are commonly tackled by separate paradigms: pixel-based regression for reconstruction, and latent diffusion for generation. Recent works attempt to unify them in latent space, but with no…

3D ReconstructionScene Generation

Improving Human Image Animation via Semantic Representation Alignment

2026-05-11 · Chang Liu, Mengting Chen, Yixuan Huang, Haoning Wu 외 arxiv

The field of image-to-video generation has made remarkable progress. However, challenges such as human limb twisting and facial distortion persist, especially when generating long videos or modeling intensive motions. Ex…

Depth EstimationVideo GenerationFace Recognition