paper-with-me

홈 › Papers

Efficient and Information-Preserving Future Frame Prediction and Beyond

2020-05-01 · ICLR 2020 1 · Wei Yu, Yichao Lu, Steve Easterbrook, Sanja Fidler

Applying resolution-preserving blocks is a common practice to maximize information preservation in video prediction, yet their high memory consumption greatly limits their application scenarios. We propose CrevNet, a Conditionally Reversible Network that uses reversible architectures to build a bijective two-way autoencoder and its complementary recurrent predictor. Our model enjoys the theoretically guaranteed property of no information loss during the feature extraction, much lower memory consumption and computational efficiency. The lightweight nature of our model enables us to incorporate 3D convolutions without concern of memory bottleneck, enhancing the model's ability to capture both short-term and long-term temporal dependencies. Our proposed approach achieves state-of-the-art results on Moving MNIST, Traffic4cast and KITTI datasets. We further demonstrate the transferability of our self-supervised learning method by exploiting its learnt features for object detection on KITTI. Our competitive results indicate the potential of using CrevNet as a generative pre-training strategy to guide downstream tasks.

📄 PDF Abstract BibTeX

Code (1)

rrxi/CrevNet pytorch

Tasks

Computational Efficiencyobject-detectionObject DetectionPredictionSelf-Supervised LearningVideo Prediction

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries

2025-10-16 · Divyat Mahajan, Sachin Goyal, Badr Youbi Idrissi, Mohammad Pezeshki 외 arxiv

Next-token prediction (NTP) has driven the success of large language models (LLMs), but it struggles with long-horizon reasoning, planning, and creative writing, with these limitations largely attributed to teacher-force…

TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction

2026-07-30 · Lei Jin, Yiding Ma, Xin Zhang, Chen Gao 외 arxiv

World Action Models (WAMs) combine future-state prediction with robot action generation, but existing approaches largely rely on visual futures. Visual prediction captures scene structure and object motion, yet provides …

ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models

2025-05-26 · Yachuan Liu, Xiaochun Wei, Lin Shi, Xinnuo Li 외

Large language models (LLMs) face significant challenges in ex-ante reasoning, where analysis, inference, or predictions must be made without access to information from future events. Even with explicit prompts enforcing…

PredictionQuestion AnsweringStock Prediction

DISTINQT: A Distributed Privacy Aware Learning Framework for QoS Prediction for Future Mobile and Wireless Networks

2024-01-15 · Nikolaos Koursioumpas, Lina Magoula, Ioannis Stavrakakis, Nancy Alonistioti 외

Beyond 5G and 6G networks are expected to support new and challenging use cases and applications that depend on a certain level of Quality of Service (QoS) to operate smoothly. Predicting the QoS in a timely manner is of…

Federated LearningPredictionPrivacy Preserving

STAU: A SpatioTemporal-Aware Unit for Video Prediction and Beyond

2022-04-20 · Zheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma 외

Video prediction aims to predict future frames by modeling the complex spatiotemporal dynamics in videos. However, most of the existing methods only model the temporal information and the spatial information for videos i…

Action Recognitionobject-detectionObject DetectionPrediction+1