paper-with-me

홈 › Papers

DriveFix: Spatio-Temporally Coherent Driving Scene Restoration

2026-03-17 · Heyu Si, Brandon James Denis, Muyang Sun, Dragos Datcu, Yaoru Li, Xin Jin, Ruiju Fu, Yuliia Tatarinova, Federico Landi, Jie Song, Mingli Song, Qi Guo arxiv

Recent advancements in 4D scene reconstruction, particularly those leveraging diffusion priors, have shown promise for novel view synthesis in autonomous driving. However, these methods often process frames independently or in a view-by-view manner, leading to a critical lack of spatio-temporal synergy. This results in spatial misalignment across cameras and temporal drift in sequences. We propose DriveFix, a novel multi-view restoration framework that ensures spatio-temporal coherence for driving scenes. Our approach employs an interleaved diffusion transformer architecture with specialized blocks to explicitly model both temporal dependencies and cross-camera spatial consistency. By conditioning the generation on historical context and integrating geometry-aware training losses, DriveFix enforces that the restored views adhere to a unified 3D geometry. This enables the consistent propagation of high-fidelity textures and significantly reduces artifacts. Extensive evaluations on the Waymo, nuScenes, and PandaSet datasets demonstrate that DriveFix achieves state-of-the-art performance in both reconstruction and novel view synthesis, marking a substantial step toward robust 4D world modeling for real-world deployment.

📄 PDF Abstract BibTeX arXiv:2603.16306

Code (0)

등록된 구현이 없습니다.

Tasks

Novel View SynthesisAutonomous Driving

Similar Papers 제목 키워드 기반

Unified Spatio-Temporal Tri-Perspective View Representation for 3D Semantic Occupancy Prediction

2024-01-24 · Sathira Silva, Savindu Bhashitha Wannigama, Gihan Jayatilaka, Muhammad Haris Khan 외

Holistic understanding and reasoning in 3D scenes play a vital role in the success of autonomous driving systems. The evolution of 3D semantic occupancy prediction as a pretraining task for autonomous driving and robotic…

3D Semantic Occupancy PredictionAutonomous DrivingComputational Efficiency

CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving

2025-08-31 · Pei Liu, Qingtian Ning, Xinyan Lu, Haipeng Liu 외 arxiv

The pursuit of autonomous agents capable of temporally coherent planning is hindered by a fundamental flaw in current vision-language models (VLMs): they lack cognitive inertia. Operating on isolated snapshots, these mod…

Knowledge DistillationAutonomous Driving

InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving

2026-06-30 · Xiaoyu Ye, Leheng Li, Xinyu Ji, Yingjie Cai 외 arxiv

Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving community. In this paper, we introduce InfiniVerse, a unified pipeline for l…

Autonomous DrivingScene Generation

4D Temporally Coherent Light-field Video

2018-04-30 · Armin Mustafa, Marco Volino, Jean-yves Guillemaut, Adrian Hilton

Light-field video has recently been used in virtual and augmented reality applications to increase realism and immersion. However, existing light-field methods are generally limited to static scenes due to the requiremen…

Scene Flow Estimation

Universal Embeddings for Spatio-Temporal Tagging of Self-Driving Logs

2020-11-12 · Sean Segal, Eric Kee, Wenjie Luo, Abbas Sadat 외

In this paper, we tackle the problem of spatio-temporal tagging of self-driving scenes from raw sensor data. Our approach learns a universal embedding for all tags, enabling efficient tagging of many attributes and faste…

BlockingTAGTemporal Tagging