paper-with-me

Papers

LongStream: Long-Sequence Streaming Autoregressive Visual Geometry

2026-02-13 · Chong Cheng, Xianda Chen, Tao Xie, Wei Yin, Weiqiang Ren, Qian Zhang, Xiaoyang Guo, Hao Wang arxiv

Long-sequence streaming 3D reconstruction remains a significant open challenge. Existing autoregressive models often fail when processing long sequences because they anchor poses to the first frame, leading to attention decay, scale drift, and extrapolation errors. We introduce LongStream, a novel gauge-decoupled streaming visual geometry model for metric-scale scene reconstruction across thousands of frames under a strictly online, future-invisible setting. Our approach is threefold. First, we discard the first-frame anchor and predict keyframe-relative poses. This reformulates long-range extrapolation into a constant-difficulty local task. Second, we introduce orthogonal scale learning. This method fully disentangles geometry from scale estimation to suppress drift. Finally, we identify attention bias issues in Transformers, including attention-sink reliance and long-term KV-cache saturation. We propose cache-consistent training combined with periodic cache refresh. This approach suppresses attention biases and contamination over ultra-long sequences and reduces the gap between training and inference. Experiments show that LongStream achieves state-of-the-art performance, enabling stable, metric-scale reconstruction over kilometer-scale sequences at 18 FPS. Project Page: https://3dagentworld.github.io/longstream/

📄 PDF Abstract BibTeX arXiv:2602.13172

Code (0)

등록된 구현이 없습니다.

Tasks

3D Reconstruction

Similar Papers 제목 키워드 기반

Rethinking the State Update Gate for Long-Sequence Recurrent 3D Reconstruction

2026-05-16 · Kejun Ren, Lei Jin, Tianxin Huang, Lianming Xu 외 arxiv

Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TTT3R-style per-token gates across five benchmarks and discover a structura…

3D ReconstructionPose Estimation

AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising

2026-03-15 · Liyuan Cui, Wentao Hu, Wenyuan Zhang, Zesong Yang 외 arxiv

Real-time talking avatar generation requires low latency and minute-level temporal stability. Autoregressive (AR) forcing enables streaming inference but suffers from exposure bias, which causes errors to accumulate and …

InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution

2025-10-01 · Ziqing Zhang, Kai Liu, Zheng Chen, Xi Li 외 arxiv

Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent challenges when processing long sequences: (1) inefficiency due to the he…

Video Super-Resolution

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars

2026-06-22 · Quanyue Song, Yishan He, Yanfei Zhang, Shihao Cheng 외 arxiv

Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to maintain visual temporal consistency and fail to explicitly perceive us…

Video Generation

SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads

2026-02-07 · Tan Yu, Qian Qiao, Le Shen, Ke Zhou 외 arxiv

Achieving a balance between high-fidelity visual quality and low-latency streaming remains a formidable challenge in audio-driven portrait generation. Existing large-scale models often suffer from prohibitive computation…

Video Generation