paper-with-me

Papers

EmerNeRF: Emergent Spatial-Temporal Scene Decomposition via Self-Supervision

2023-11-03 · Jiawei Yang, Boris Ivanovic, Or Litany, Xinshuo Weng, Seung Wook Kim, Boyi Li, Tong Che, Danfei Xu, Sanja Fidler, Marco Pavone, Yue Wang

We present EmerNeRF, a simple yet powerful approach for learning spatial-temporal representations of dynamic driving scenes. Grounded in neural fields, EmerNeRF simultaneously captures scene geometry, appearance, motion, and semantics via self-bootstrapping. EmerNeRF hinges upon two core components: First, it stratifies scenes into static and dynamic fields. This decomposition emerges purely from self-supervision, enabling our model to learn from general, in-the-wild data sources. Second, EmerNeRF parameterizes an induced flow field from the dynamic field and uses this flow field to further aggregate multi-frame features, amplifying the rendering precision of dynamic objects. Coupling these three fields (static, dynamic, and flow) enables EmerNeRF to represent highly-dynamic scenes self-sufficiently, without relying on ground truth object annotations or pre-trained models for dynamic object segmentation or optical flow estimation. Our method achieves state-of-the-art performance in sensor simulation, significantly outperforming previous methods when reconstructing static (+2.93 PSNR) and dynamic (+3.70 PSNR) scenes. In addition, to bolster EmerNeRF's semantic generalization, we lift 2D visual foundation model features into 4D space-time and address a general positional bias in modern Transformers, significantly boosting 3D perception performance (e.g., 37.50% relative improvement in occupancy prediction accuracy on average). Finally, we construct a diverse and challenging 120-sequence dataset to benchmark neural fields under extreme and highly-dynamic settings.

📄 PDF Abstract BibTeX arXiv:2311.02077

Code (1)

nvlabs/emernerf pytorch

Tasks

Optical Flow EstimationSemantic Segmentation

Similar Papers 제목 키워드 기반

Hybrid Congestion Classification Framework Using Flow-Guided Attention and Empirical Mode Decomposition

2026-05-06 · Eugene Kofi Okrah Denteh, Blessing Agyei Kyem, Joshua Kofi Asamoah, Armstrong Aboah arxiv

Accurate traffic congestion classification requires models that jointly capture roadway scene context and non-stationary traffic motion, yet most prior work treats these requirements in isolation. Vision-based methods of…

UnIRe: Unsupervised Instance Decomposition for Dynamic Urban Scene Reconstruction

2025-04-01 · Yunxuan Mao, Rong Xiong, Yue Wang, Yiyi Liao

Reconstructing and decomposing dynamic urban scenes is crucial for autonomous driving, urban planning, and scene editing. However, existing methods fail to perform instance-aware decomposition without manual annotations,…

3DGSAutonomous Driving

Polarimetric Spatio-Temporal Light Transport Probing

2021-05-25 · Seung-Hwan Baek, Felix Heide

Light emitted from a source into a scene can undergo complex interactions with scene surfaces of different material types before being reflected. During this transport, every surface reflection is encoded in the properti…

MetamerismScene Understanding

Grouped Spatial-Temporal Aggregation for Efficient Action Recognition

2019-09-28 · ICCV 2019 10 · Chenxu Luo, Alan Yuille

Temporal reasoning is an important aspect of video analysis. 3D CNN shows good performance by exploring spatial-temporal features jointly in an unconstrained way, but it also increases the computational cost a lot. Previ…

Action Recognition

Decomposition, Compression, and Synthesis (DCS)-based Video Coding: A Neural Exploration via Resolution-Adaptive Learning

2020-12-01 · Ming Lu, Tong Chen, Dandan Ding, Fengqing Zhu 외

Inspired by the facts that retinal cells actually segregate the visual scene into different attributes (e.g., spatial details, temporal motion) for respective neuronal processing, we propose to first decompose the input …

Motion CompensationSuper-ResolutionVideo CompressionVideo Reconstruction