paper-with-me

홈 › Papers

Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction

2024-10-24 · Junyi Chen, Di Huang, Weicai Ye, Wanli Ouyang, Tong He

Spatial intelligence is the ability of a machine to perceive, reason, and act in three dimensions within space and time. Recent advancements in large-scale auto-regressive models have demonstrated remarkable capabilities across various reasoning tasks. However, these models often struggle with fundamental aspects of spatial reasoning, particularly in answering questions like "Where am I?" and "What will I see?". While some attempts have been done, existing approaches typically treat them as separate tasks, failing to capture their interconnected nature. In this paper, we present Generative Spatial Transformer (GST), a novel auto-regressive framework that jointly addresses spatial localization and view prediction. Our model simultaneously estimates the camera pose from a single image and predicts the view from a new camera pose, effectively bridging the gap between spatial awareness and visual prediction. The proposed innovative camera tokenization method enables the model to learn the joint distribution of 2D projections and their corresponding spatial perspectives in an auto-regressive manner. This unified training paradigm demonstrates that joint optimization of pose estimation and novel view synthesis leads to improved performance in both tasks, for the first time, highlighting the inherent relationship between spatial awareness and visual prediction.

📄 PDF Abstract BibTeX arXiv:2410.18962

Code (0)

등록된 구현이 없습니다.

Tasks

Novel View SynthesisPose EstimationPredictionSpatial Reasoning

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Multi-Head Attention 설명 없음
Adam 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…

Similar Papers 제목 키워드 기반

From "What" to "How": Constrained Reasoning for Autoregressive Image Generation

2026-03-03 · Ruxue Yan, Xubo Liu, Wenya Guo, Zhengkun Zhang 외 arxiv

Autoregressive image generation has seen recent improvements with the introduction of chain-of-thought and reinforcement learning. However, current methods merely specify "What" details to depict by rewriting the input p…

Reinforcement LearningImage Generation

FIction: 4D Future Interaction Prediction from Video

2024-12-01 · CVPR 2025 1 · Kumar Ashutosh, Georgios Pavlakos, Kristen Grauman

Anticipating how a person will interact with objects in an environment is essential for activity understanding, but existing methods are limited to the 2D space of video frames-capturing physically ungrounded predictions…

Prediction

Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech Recognition

2020-05-16 · Zhengkun Tian, Jiangyan Yi, Jian-Hua Tao, Ye Bai 외

Non-autoregressive transformer models have achieved extremely fast inference speed and comparable performance with autoregressive sequence-to-sequence models in neural machine translation. Most of the non-autoregressive …

Machine Translationspeech-recognitionSpeech RecognitionTranslation

Beyond Spatial Auto-Regressive Models: Predicting Housing Prices with Satellite Imagery

2016-10-16 · Archith J. Bency, Swati Rallapalli, Raghu K. Ganti, Mudhakar Srivatsa 외

When modeling geo-spatial data, it is critical to capture spatial correlations for achieving high accuracy. Spatial Auto-Regression (SAR) is a common tool used to model such data, where the spatial contiguity matrix (W) …

Seg-VAR: Image Segmentation with Visual Autoregressive Modeling

2025-11-16 · Rongkun Zheng, Lu Qi, Xi Chen, Yi Wang 외 arxiv

While visual autoregressive modeling (VAR) strategies have shed light on image generation with the autoregressive models, their potential for segmentation, a task that requires precise low-level spatial perception, remai…

Image SegmentationImage Generation