paper-with-me

홈 › Papers

Predicting Long-horizon Futures by Conditioning on Geometry and Time

2024-04-17 · Tarasha Khurana, Deva Ramanan

Our work explores the task of generating future sensor observations conditioned on the past. We are motivated by `predictive coding' concepts from neuroscience as well as robotic applications such as self-driving vehicles. Predictive video modeling is challenging because the future may be multi-modal and learning at scale remains computationally expensive for video processing. To address both challenges, our key insight is to leverage the large-scale pretraining of image diffusion models which can handle multi-modality. We repurpose image models for video prediction by conditioning on new frame timestamps. Such models can be trained with videos of both static and dynamic scenes. To allow them to be trained with modestly-sized datasets, we introduce invariances by factoring out illumination and texture by forcing the model to predict (pseudo) depth, readily obtained for in-the-wild videos via off-the-shelf monocular depth networks. In fact, we show that simply modifying networks to predict grayscale pixels already improves the accuracy of video prediction. Given the extra controllability with timestamp conditioning, we propose sampling schedules that work better than the traditional autoregressive and hierarchical sampling strategies. Motivated by probabilistic metrics from the object forecasting literature, we create a benchmark for video prediction on a diverse set of videos spanning indoor and outdoor scenes and a large vocabulary of objects. Our experiments illustrate the effectiveness of learning to condition on timestamps, and show the importance of predicting the future with invariant modalities.

📄 PDF Abstract BibTeX arXiv:2404.11554

Code (0)

등록된 구현이 없습니다.

Tasks

Video Prediction

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Risk Horizons: Structured Hypothesis Spaces for Longitudinal Clinical Prediction

2026-02-13 · Zhan Qu, Michael Färber arxiv

Predicting future clinical events from longitudinal electronic health records (EHRs) requires selecting plausible outcomes from a large and structured event space under sparse observations. While clinical coding systems …

EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory

2025-10-01 · Jiahao Wang, Luoxin Ye, TaiMing Lu, Junfei Xiao 외 arxiv

Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world model that bridges panoramic video genera…

3D ReconstructionVideo Generation

LLM-Guided Future Hypotheses for Horizon-Aware Exploration in Multi-Step Robot Manipulation

2026-05-28 · Mohammad Khoshnazar, Andrew Melnik, Michael Beetz arxiv

Multi-step robot manipulation requires acting under uncertainty about how the scene will evolve, making exploration and policy adaptation challenging. We study whether short-horizon, task-consistent future videos can pro…

Robot Manipulation

How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction

2026-05-19 · Sejoon Jun, Hai Nguyen-Truong, Luigi Seminara, Lorenzo Torresani arxiv

Predicting how a person's first-person view will evolve (what action will follow, what plan completes a task, whether an in-progress shot will score) is fundamentally under-specified: the same context admits many plausib…

Pose Estimation

Trading Signals In VIX Futures

2021-03-02 · M. Avellaneda, T. N. Li, A. Papanicolaou, G. Wang

We propose a new approach for trading VIX futures. We assume that the term structure of VIX futures follows a Markov model. Our trading strategy selects a position in VIX futures by maximizing the expected utility for a …

Position