paper-with-me

Papers

APT: Atomic Physical Transitions for Causal Video-Language Understanding

2026-06-17 · Shang Wu, Haoran Lu, Songling Liu, Chenwei Xu, Lie Lu, Pranav Maneriker, Fan Du, Manling Li, Zhaoran Wang, Han Liu arxiv

Physical events are not understood by their names alone, but by the causal state changes that compose them. A clip-level label such as "bounce" can be correct while hiding the process that makes the event physically valid, from support loss and contact onset to rebound and settling. To make this hidden process explicit, we introduce Atomic Physical Transitions (APTs): minimal, temporally localized state changes that bind a visible cue to an active physical mechanism and before/after dynamical regimes. An APT chain represents a video as an ordered causal transition sequence rather than a single aggregate event label: event labels tell what happened; APT chains explain why it happened. To make APTs learnable by VLMs, we construct mixed-source APT data from human annotations and simulator ground truth, covering 14 transition types across contact, gravity, friction, and rotation/stability, with 27,303 timed instances over 1,246 trials. Using this data, we find that current VLMs miss transition-level physics, with zero-shot recall at most 14% and errors dominated by missed transitions. Direct fine-tuning on APT chains improves transition detection but causes event-level forgetting, indicating that the model learns a specialized answer format rather than a reusable physical representation. We therefore propose APT-Tune, a parameter-efficient recipe that teaches VLMs to use causal transitions without forgetting how to answer video questions. It combines image-pad-aware supervision, format-conditional co-training, and mechanism-conditioned domain-to-type decoding to make APT learning format-robust and physically grounded. With only 11 M LoRA parameters on Qwen3-VL-2B, APT-Tune substantially improves APT recall while also improving event-level video transfer. These results show that APTs are not a new answer format, but a human-aligned causal supervision signal for physical video understanding.

📄 PDF Abstract BibTeX arXiv:2606.18586

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CausalMotion: Structured Physical Reasoning as Keyframe and Trajectory Guidance for Training-Free Video Generation

2026-06-12 · Sihan Zhuang, Xinyuan Chen, Tianfan Xue, Yaohui Wang arxiv

Recent advances in diffusion-based video generation have significantly improved visual quality and short-term temporal coherence. However, existing methods still struggle to produce videos with physically consistent and …

Video Generation

PhiZero: A World Model Built Around Physical Language

2026-07-30 · Shuyao Shang, Yuqi Wang, Ruopeng Gao, Xu Chen 외 arxiv

We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel …

Chain of Event-Centric Causal Thought for Physically Plausible Video Generation

2026-03-10 · Zixuan Wang, Yixin Hu, Haolan Wang, Feng Chen 외 arxiv

Physically Plausible Video Generation (PPVG) has emerged as a promising avenue for modeling real-world physical phenomena. PPVG requires an understanding of commonsense knowledge, which remains a challenge for video diff…

Video Generation

X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling

2026-05-24 · Baolu Li, Jingyu Qian, Rui Guo, Yilun Chen 외 arxiv

Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and generalizable planning. Predictive world modeling enables VLA to internaliz…

CLEVRER-Humans: Describing Physical and Causal Events the Human Way

2023-10-05 · Jiayuan Mao, Xuelin Yang, Xikun Zhang, Noah D. Goodman 외

Building machines that can reason about physical events and their causal relationships is crucial for flexible interaction with the physical world. However, most existing physical and causal reasoning benchmarks are excl…

Causal JudgmentData AugmentationDiversityQuestion Answering