paper-with-me

홈 › Papers

MORPHOS: Autoregressive 4D Generation with Temporal Structured Latents

2026-06-01 · Minkyung Kwon, Jinhyeok Choi, Youngjin Shin, Jaeyeong Kim, JongMin Lee, Seungryong Kim arxiv

We present MORPHOS, a novel autoregressive framework that generates dynamic 3D assets from videos across diverse representations, including meshes, 3D Gaussians, and radiance fields. Existing methods are typically limited to a single representation, struggle to model topological changes, or fail to maintain temporal consistency over long videos. To address these limitations, we introduce the Temporal Structured Latents (T-SLAT), a unified 4D representation that jointly encodes geometry and appearance along the temporal dimension. Leveraging T-SLAT, MORPHOS autoregressively generates dynamic 3D assets via causal attention, conditioning each frame on its preceding history to ensure temporal consistency while handling evolving topologies. We also propose a temporal-structural augmentation to mitigate error accumulation in autoregressive generation. MORPHOS achieves state-of-the-art performance in appearance and competitive results in geometry across multiple benchmarks, demonstrating superior generalization across various representations and robustness in long-horizon generation.

📄 PDF Abstract BibTeX arXiv:2606.02491

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SAGA: Stable Acceleration Guidance for Autoregressive Video Generation

2026-07-09 · Thanh-Nhan Vo, Trong-Thuan Nguyen, Trung-Hoang Le, Tam V. Nguyen 외 arxiv

Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context can amplify temporal errors, resulting in flickering, motion jitter…

Video Generation

UniTemp: Unlocking Video Generation in Any Temporal Order via Bidirectional Distillation

2026-06-17 · Lin Zhang, Sicheng Mo, Zefan Cai, Jinhong Lin 외 arxiv

Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods are restricted to forward temporal gener…

Story GenerationVideo Generation

MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent Space

2025-03-19 · Lixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan 외

This paper addresses the challenge of text-conditioned streaming motion generation, which requires us to predict the next-step human pose based on variable-length historical motions and incoming texts. Existing methods s…

Motion Generation

Semantic Image Synthesis with Semantically Coupled VQ-Model

2022-09-06 · Stephan Alaniz, Thomas Hummel, Zeynep Akata

Semantic image synthesis enables control over unconditional image generation by allowing guidance on what is being generated. We conditionally synthesize the latent space from a vector quantized model (VQ-model) pre-trai…

Image GenerationUnconditional Image Generation

Self Gradient Forcing: Native Long Video Extrapolation

2026-07-22 · Junhao Zhuang, Shiyi Zhang, Yuxuan Bian, Yaowei Li 외 arxiv

Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure…

Video Generation