paper-with-me

홈 › Papers

VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement

2026-01-20 · Tiancheng Fang, Bowen Pan, Lingxi Chen, Jiangjing Lyu, Chengfei Lyu, Chaoyue Niu, Fan Wu arxiv

We propose VIAFormer, a Voxel-Image Alignment Transformer model designed for Multi-view Conditioned Voxel Refinement--the task of repairing incomplete noisy voxels using calibrated multi-view images as guidance. Its effectiveness stems from a synergistic design: an Image Index that provides explicit 3D spatial grounding for 2D image tokens, a Correctional Flow objective that learns a direct voxel-refinement trajectory, and a Hybrid Stream Transformer that enables robust cross-modal fusion. Experiments show that VIAFormer establishes a new state of the art in correcting both severe synthetic corruptions and realistic artifacts on the voxel shape obtained from powerful Vision Foundation Models. Beyond benchmarking, we demonstrate VIAFormer as a practical and reliable bridge in real-world 3D creation pipelines, paving the way for voxel-based methods to thrive in large-model, big-data wave.

📄 PDF Abstract BibTeX arXiv:2601.13664

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation

2026-06-23 · Haorui Ji, Weizhe Liu, Hongdong Li, Hengkai Guo arxiv

Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation, yet current methods struggle to preserve high-frequency visual details of input images due to two str…

Representation Learning

Multi-Resolution Alignment for Voxel Sparsity in Camera-Based 3D Semantic Scene Completion

2026-02-03 · Zhiwen Yang, Yuxin Peng arxiv

Camera-based 3D semantic scene completion (SSC) offers a cost-effective solution for assessing the geometric occupancy and semantic labels of each voxel in the surrounding 3D scene with image inputs, providing a voxel-le…

3D Semantic Scene CompletionAutonomous Driving

Meta-Learning an In-Context Transformer Model of Human Higher Visual Cortex

2025-05-21 · Muquan Yu, Mu Nan, Hossein Adeli, Jacob S. Prince 외

Understanding functional representations within higher visual cortex is a fundamental question in computational neuroscience. While artificial neural networks pretrained on large-scale datasets exhibit striking represent…

In-Context LearningInductive BiasMeta-LearningNatural Language Queries

Deep-Learning-based Fast and Accurate 3D CT Deformable Image Registration in Lung Cancer

2023-04-21 · Yuzhen Ding, Hongying Feng, Yunze Yang, Jason Holmes 외

Purpose: In some proton therapy facilities, patient alignment relies on two 2D orthogonal kV images, taken at fixed, oblique angles, as no 3D on-the-bed imaging is available. The visibility of the tumor in kV images is l…

AnatomyImage Registration

VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI

2025-09-10 · Chenqian Le, Yilin Zhao, Nikasadat Emami, Kushagra Yadav 외 arxiv

Recent advances in fMRI-based visual decoding have enabled compelling reconstructions of perceived images. However, most approaches rely on subject-specific training, limiting scalability and practical deployment. We int…