paper-with-me

홈 › Papers

Efficient 3D Object Reconstruction using Visual Transformers

2023-02-16 · Rohan Agarwal, Wei Zhou, Xiaofeng Wu, Yuhan Li

Reconstructing a 3D object from a 2D image is a well-researched vision problem, with many kinds of deep learning techniques having been tried. Most commonly, 3D convolutional approaches are used, though previous work has shown state-of-the-art methods using 2D convolutions that are also significantly more efficient to train. With the recent rise of transformers for vision tasks, often outperforming convolutional methods, along with some earlier attempts to use transformers for 3D object reconstruction, we set out to use visual transformers in place of convolutions in existing efficient, high-performing techniques for 3D object reconstruction in order to achieve superior results on the task. Using a transformer-based encoder and decoder to predict 3D structure from 2D images, we achieve accuracy similar or superior to the baseline approach. This study serves as evidence for the potential of visual transformers in the task of 3D object reconstruction.

📄 PDF Abstract BibTeX arXiv:2302.08474

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object ReconstructionDecoderObjectObject Reconstruction

Similar Papers 제목 키워드 기반

3D-RETR: End-to-End Single and Multi-View 3D Reconstruction with Transformers

2021-10-17 · Zai Shi, Zhao Meng, Yiran Xing, Yunpu Ma 외

3D reconstruction aims to reconstruct 3D objects from 2D views. Previous works for 3D reconstruction mainly focus on feature matching between views or using CNNs as backbones. Recently, Transformers have been shown effec…

3D ReconstructionDecoderMulti-View 3D Reconstruction

Visual Attention Emerges from Recurrent Sparse Reconstruction

2022-04-23 · Baifeng Shi, Yale Song, Neel Joshi, Trevor Darrell 외

Visual attention helps achieve robust perception under noise, corruption, and distribution shifts in human vision, which are areas where modern neural networks still fall short. We present VARS, Visual Attention from Rec…

Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers

2025-11-18 · Yutian Chen, Yuheng Qiu, Ruogu Li, Ali Agha 외 arxiv

We propose Confidence-Guided Token Merging (Co-Me), an acceleration mechanism for visual geometric transformers without retraining or finetuning the base model. Co-Me distilled a light-weight confidence predictor to rank…

TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention

2026-05-14 · David Huang, Guile Wu, Chengjie Huang, Bingbing Liu 외 arxiv

Recent feed-forward 3D reconstruction methods, such as visual geometry transformers, have substantially advanced the traditional per-scene optimization paradigm by enabling effective multi-view reconstruction in a single…

Multi-View 3D ReconstructionComputational Efficiency

RIS-Assisted 3D Spherical Splatting for Object Composition Visualization using Detection Transformers

2025-11-04 · Anastasios T. Sotiropoulos, Stavros Tsimpoukis, Dimitrios Tyrovolas, Sotiris Ioannidis 외 arxiv

The pursuit of immersive and structurally aware multimedia experiences has intensified interest in sensing modalities that reconstruct objects beyond the limits of visible light. Conventional optical pipelines degrade un…