paper-with-me

Papers

Spatio-Temporal Outdoor Lighting Aggregation on Image Sequences using Transformer Networks

2022-02-18 · Haebom Lee, Christian Homeyer, Robert Herzog, Jan Rexilius, Carsten Rother

In this work, we focus on outdoor lighting estimation by aggregating individual noisy estimates from images, exploiting the rich image information from wide-angle cameras and/or temporal image sequences. Photographs inherently encode information about the scene's lighting in the form of shading and shadows. Recovering the lighting is an inverse rendering problem and as that ill-posed. Recent work based on deep neural networks has shown promising results for single image lighting estimation, but suffers from robustness. We tackle this problem by combining lighting estimates from several image views sampled in the angular and temporal domain of an image sequence. For this task, we introduce a transformer architecture that is trained in an end-2-end fashion without any statistical post-processing as required by previous work. Thereby, we propose a positional encoding that takes into account the camera calibration and ego-motion estimation to globally register the individual estimates when computing attention between visual words. We show that our method leads to improved lighting estimation while requiring less hyper-parameters compared to the state-of-the-art.

📄 PDF Abstract BibTeX arXiv:2202.09206

Code (0)

등록된 구현이 없습니다.

Tasks

Camera CalibrationInverse RenderingLighting EstimationMotion Estimation

Similar Papers 제목 키워드 기반

Learning to Factorize and Relight a City

2020-08-06 · ECCV 2020 8 · Andrew Liu, Shiry Ginosar, Tinghui Zhou, Alexei A. Efros 외

We propose a learning-based framework for disentangling outdoor scenes into temporally-varying illumination and permanent scene factors. Inspired by the classic intrinsic image decomposition, our learning signal builds u…

Intrinsic Image Decomposition

Lighting in Motion: Spatiotemporal HDR Lighting Estimation

2025-12-15 · Christophe Bolduc, Julien Philip, Li Ma, Mingming He 외 arxiv

We present Lighting in Motion (LiMo), a diffusion-based approach to spatiotemporal lighting estimation. LiMo targets both realistic high-frequency detail prediction and accurate illuminance estimation. To account for bot…

Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition

2026-05-22 · Yerim So, Jiyeong Kim, Jiwon Yoon, Dongbo Min arxiv

Recent Open-Vocabulary Action Recognition (OVAR) methods typically aggregate visual features into a global representation before computing text alignment, a process that obscures local patch information and fine-grained …

Action Recognition

Spatiotemporally Consistent Indoor Lighting Estimation with Diffusion Priors

2025-08-11 · Mutian Tong, Rundi Wu, Changxi Zheng arxiv

Indoor lighting estimation from a single image or video remains a challenge due to its highly ill-posed nature, especially when the lighting condition of the scene varies spatially and temporally. We propose a method tha…

Zero-shot Generalization

Learning Spatial and Spatio-Temporal Pixel Aggregations for Image and Video Denoising

2021-01-26 · Xiangyu Xu, Muchen Li, Wenxiu Sun, Ming-Hsuan Yang

Existing denoising methods typically restore clear results by aggregating pixels from the noisy input. Instead of relying on hand-crafted aggregation schemes, we propose to explicitly learn this process with deep neural …

DenoisingImage DenoisingVideo Denoising