paper-with-me

홈 › Papers

Priors are Powerful: Improving a Transformer for Multi-camera 3D Detection with 2D Priors

2023-01-31 · Di Feng, Francesco Ferroni

Transfomer-based approaches advance the recent development of multi-camera 3D detection both in academia and industry. In a vanilla transformer architecture, queries are randomly initialised and optimised for the whole dataset, without considering the differences among input frames. In this work, we propose to leverage the predictions from an image backbone, which is often highly optimised for 2D tasks, as priors to the transformer part of a 3D detection network. The method works by (1). augmenting image feature maps with 2D priors, (2). sampling query locations via ray-casting along 2D box centroids, as well as (3). initialising query features with object-level image features. Experimental results shows that 2D priors not only help the model converge faster, but also largely improve the baseline approach by up to 12% in terms of average precision.

📄 PDF Abstract BibTeX arXiv:2301.13592

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Divide and Conquer: Improving Multi-Camera 3D Perception with 2D Semantic-Depth Priors and Input-Dependent Queries

2024-08-13 · Qi Song, Qingyong Hu, Chi Zhang, Yongquan Chen 외

3D perception tasks, such as 3D object detection and Bird's-Eye-View (BEV) segmentation using multi-camera images, have drawn significant attention recently. Despite the fact that accurately estimating both semantic and …

3D Object DetectionBEV SegmentationObjectObject Categorization+3

Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors

2025-03-21 · CVPR 2025 1 · Wonbong Jang, Philippe Weinzaepfel, Vincent Leroy, Lourdes Agapito 외

We present Pow3r, a novel large 3D vision regression model that is highly versatile in the input modalities it accepts. Unlike previous feed-forward models that lack any mechanism to exploit known camera or scene priors …

3D ReconstructionDepth CompletionDepth EstimationDepth Prediction+2

Scene Coordinate Reconstruction Priors

2025-10-14 · Wenjing Bian, Axel Barroso-Laguna, Tommaso Cavallari, Victor Adrian Prisacariu 외 arxiv

Scene coordinate regression (SCR) models have proven to be powerful implicit scene representations for 3D vision, enabling visual relocalization and structure-from-motion. SCR models are trained specifically for one scen…

Novel View SynthesisPoint Clouds

VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection

2026-03-01 · Yang Cao, Feize Wu, Dave Zhenyu Chen, Yingji Zhong 외 arxiv

Current multi-view indoor 3D object detectors rely on sensor geometry that is costly to obtain (i.e., precisely calibrated multi-view camera poses) to fuse multi-view information into a global scene representation, limit…

3D Object Detection

A Light Touch Approach to Teaching Transformers Multi-view Geometry

2022-11-28 · CVPR 2023 1 · Yash Bhalgat, Joao F. Henriques, Andrew Zisserman

Transformers are powerful visual learners, in large part due to their conspicuous lack of manually-specified priors. This flexibility can be problematic in tasks that involve multiple-view geometry, due to the near-infin…

Retrieval