paper-with-me

홈 › Papers

UniQuery4R: Unified 4D Scene Reconstruction from a Single Query

2026-08-18 · Tiancheng Chen, Sheng Tang, Wenhua Jin, Weiqi Zhang, Juntong Fang, Junsheng Zhou, Zesong Li arxiv

Reconstructing dynamic 4D scenes requires jointly estimating correspondence, geometry, object motion, and camera motion. Existing feed-forward methods typically predict dense task-specific maps or independently process source-target pairs, leading to unnecessary computation for sparse queries and limited feature reuse across different frame pairs. We present UniQuery4R, a query-conditioned framework that encodes a multi-frame clip once and selects the source view, target view, and continuous source-image coordinate only at decoding time via source-to-target cross-attention. Each query jointly predicts target correspondence, target-time 3D position, and scene flow, along with source depth, while camera parameters are estimated per view. This design allows the encoded clip to be reused across arbitrary source-target selections and supports both sparse inference and dense reconstruction through batched queries, without learned temporal embeddings tied to a fixed clip length. We further introduce a direction-magnitude parameterization of scene flow with separate supervision for moving and static points. Among the evaluated methods, UniQuery4R achieves the best macro-average results on WorldTrack for both scene-flow estimation and dynamic-point reconstruction.

📄 PDF Abstract BibTeX arXiv:2608.17283

Code (2)

Tavish9/awesome-daily-AI-arxiv ★ 113
cakerdsp/geometry-vision-daily ★ 2

Similar Papers 제목 키워드 기반

UniQueR: Unified Query-based Feedforward 3D Reconstruction

2026-03-24 · Chensheng Peng, Quentin Herau, Jiezhi Yang, Yichen Xie 외 arxiv

We present UniQueR, a unified query-based feedforward framework for efficient and accurate 3D reconstruction from unposed images. Existing feedforward models such as DUSt3R, VGGT, and AnySplat typically predict per-pixel…

3D Reconstruction

PanopticQuery: Unified Query-Time Reasoning for 4D Scenes

2026-04-07 · Ruilin Tang, Yang Zhou, Zhong Ye, Wenxi Liu 외 arxiv

Understanding dynamic 4D environments through natural language queries requires not only accurate scene reconstruction but also robust semantic grounding across space, time, and viewpoints. While recent methods using neu…

Natural Language QueriesDynamic Reconstruction

4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere

2026-02-10 · Yihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan 외 arxiv

We present 4RC, a unified feed-forward framework for 4D reconstruction from monocular videos. Unlike existing approaches that typically decouple motion from geometry or produce limited 4D attributes such as sparse trajec…

PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction

2026-03-06 · Xiang Zhang, Sohyun Yoo, Hongrui Wu, Chuan Li 외 arxiv

We introduce PixARMesh, a method to autoregressively reconstruct complete 3D indoor scene meshes directly from a single RGB image. Unlike prior methods that rely on implicit signed distance fields and post-hoc layout opt…

Spatial Reasoning

Segment then Splat: A Unified Approach for 3D Open-Vocabulary Segmentation based on Gaussian Splatting

2025-03-28 · Yiren Lu, Yunlai Zhou, Yiran Qiao, Chaoda Song 외

Open-vocabulary querying in 3D space is crucial for enabling more intelligent perception in applications such as robotics, autonomous systems, and augmented reality. However, most existing methods rely on 2D pixel-level …

3D Object RetrievalObjectSegmentation