paper-with-me

홈 › Papers

PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop

2025-03-12 · Chenyu Li, Oscar Michel, Xichen Pan, Sainan Liu, Mike Roberts, Saining Xie

Large-scale pre-trained video generation models excel in content creation but are not reliable as physically accurate world simulators out of the box. This work studies the process of post-training these models for accurate world modeling through the lens of the simple, yet fundamental, physics task of modeling object freefall. We show state-of-the-art video generation models struggle with this basic task, despite their visually impressive outputs. To remedy this problem, we find that fine-tuning on a relatively small amount of simulated videos is effective in inducing the dropping behavior in the model, and we can further improve results through a novel reward modeling procedure we introduce. Our study also reveals key limitations of post-training in generalization and distribution modeling. Additionally, we release a benchmark for this task that may serve as a useful diagnostic tool for tracking physical accuracy in large-scale video generative model development.

📄 PDF Abstract BibTeX arXiv:2503.09595

Code (1)

vision-x-nyu/pisa-experiments 공식 구현 pytorch

Tasks

DiagnosticVideo Generation

Similar Papers 제목 키워드 기반

PISA: Pixelwise Image Saliency by Aggregating Complementary Appearance Contrast Measures with Spatial Priors

2013-06-01 · CVPR 2013 6 · Keyang Shi, Keze Wang, Jiangbo Lu, Liang Lin

Driven by recent vision and graphics applications such as image segmentation and object recognition, assigning pixel-accurate saliency values to uniformly highlight foreground objects becomes increasingly critical. More …

Image SegmentationObject RecognitionSaliency DetectionSemantic Segmentation

Physically Informed Synchronic-adaptive Learning for Industrial Systems Modeling in Heterogeneous Media with Unavailable Time-varying Interface

2024-01-26 · Aina Wang, Pan Qin, Xi-Ming Sun

Partial differential equations (PDEs) are commonly employed to model complex industrial systems characterized by multivariable dependence. Existing physics-informed neural networks (PINNs) excel in solving PDEs in a homo…

Physics-Informed Structure Anchoring With Capture-Aware Prototype Calibration for Cross-Environment RF Fingerprinting

2026-07-06 · Fengchong Yao, Jianbing Li, Qing Liu, Qikun Liu 외 arxiv

Radio frequency fingerprint identification (RFFI) exploits transmitter-specific hardware imperfections as physicallayer identity cues for Internet of Things (IoT) devices, but deep models often degrade across acquisition…

Representation Learning

PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models

2025-03-13 · Zilu Guo, Hongbin Lin, Zhihao Yuan, Chaoda Zheng 외

3D Multimodal Large Language Models (MLLMs) have recently made substantial advancements. However, their potential remains untapped, primarily due to the limited quantity and suboptimal quality of 3D datasets. Current app…

3D Object Captioning

PISA: Pixelwise Image Saliency by Aggregating Complementary Appearance Contrast Measures with Edge-Preserving Coherence

2015-05-13 · CVPR 2013 · Keze Wang, Liang Lin, Jiangbo Lu, Chenglong Li 외

Driven by recent vision and graphics applications such as image segmentation and object recognition, computing pixel-accurate saliency values to uniformly highlight foreground objects becomes increasingly important. In t…

Image SegmentationObject RecognitionSaliency DetectionSemantic Segmentation