paper-with-me

홈 › Papers

Guiding Video Prediction with Explicit Procedural Knowledge

2024-06-26 · Patrick Takenaka, Johannes Maucher, Marco F. Huber

We propose a general way to integrate procedural knowledge of a domain into deep learning models. We apply it to the case of video prediction, building on top of object-centric deep models and show that this leads to a better performance than using data-driven models alone. We develop an architecture that facilitates latent space disentanglement in order to use the integrated procedural knowledge, and establish a setup that allows the model to learn the procedural interface in the latent space using the downstream task of video prediction. We contrast the performance to a state-of-the-art data-driven approach and show that problems where purely data-driven approaches struggle can be handled by using knowledge about the domain, providing an alternative to simply collecting more data.

📄 PDF Abstract BibTeX arXiv:2406.18220

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementPredictionVideo Prediction

Similar Papers 제목 키워드 기반

ViPro: Enabling and Controlling Video Prediction for Complex Dynamical Scenarios using Procedural Knowledge

2024-06-26 · Patrick Takenaka, Johannes Maucher, Marco F. Huber

We propose a novel architecture design for video prediction in order to utilize procedural domain knowledge directly as part of the computational graph of data-driven models. On the basis of new challenging scenarios we …

Video Prediction

Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training

2025-02-24 · Karan Samel, Nitish Sontakke, Irfan Essa

Instructional videos provide a convenient modality to learn new tasks (ex. cooking a recipe, or assembling furniture). A viewer will want to find a corresponding video that reflects both the overall task they are interes…

ViPro-2: Unsupervised State Estimation via Integrated Dynamics for Guiding Video Prediction

2025-08-08 · Patrick Takenaka, Johannes Maucher, Marco F. Huber arxiv

Predicting future video frames is a challenging task with many downstream applications. Previous work has shown that procedural knowledge enables deep models for complex dynamical settings, however their model ViPro assu…

Video Prediction

Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO

2025-11-20 · Junhao Cheng, Liang Hou, Xin Tao, Jing Liao arxiv

While language models have become impactful in many real-world applications, video generation remains largely confined to entertainment. Motivated by video's inherent capacity to demonstrate physical-world information th…

Reinforcement LearningVideo Generation

Factorizing Declarative and Procedural Knowledge in Structured, Dynamical Environments

2021-01-01 · ICLR 2021 1 · Anirudh Goyal, Alex Lamb, Phanideep Gampa, Philippe Beaudoin 외

Modeling a structured, dynamic environment like a video game requires keeping track of the objects and their states (declarative knowledge) as well as predicting how objects behave (procedural knowledge). Black-box model…

Object