paper-with-me

홈 › Papers

VideoWeave: Unlocking Geometric Consistency in Video Generation via Joint Geometry-Video Modeling

2026-06-12 · Xunzhi Xiang, Zixuan Duan, Yabo Chen, Zhengxuan Wei, Guiyu Zhang, Zixiao Gu, Zhe Gao, Haibin Huang, Chi Zhang, Qi Fan, Xuelong Li arxiv

Large-scale video diffusion models often fail to preserve 3D structure over time, causing geometric drift and implausible motion under viewpoint changes. Existing methods usually enforce geometric consistency by using explicit geometry reconstructions, such as depth maps, point clouds, or reconstructed 3D structures, to define conditions, supervision, or reward signals, making the generator sensitive to errors from upstream geometry pipelines. We propose VideoWeave, a latent-space post-training framework that uses implicit geometry-model features to constrain the generative distribution, providing a more flexible and non-rigid form of guidance that mitigates the impact of reconstruction errors from geometry models. Specifically, VideoWeave adapts these features into geometry latents and jointly models them with video latents in a shared denoising space, allowing geometry to shape the generative distribution during training. To support this process, we build GeoVid-80K, an 80K-video dataset with paired appearance and geometry representations. Experiments on text-to-video and image-to-video generation show that VideoWeave improves geometric coherence while preserving strong visual quality. VideoWeave project page at https://videoweave.github.io/

📄 PDF Abstract BibTeX arXiv:2606.14162

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationPoint Clouds

Similar Papers 제목 키워드 기반

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation

2026-06-06 · Jianhui Wei, Jie Tan, Hengchuan Zhu, Xiaotian Zhang 외 arxiv

Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video generation, a long-horizon multimodal task, remains underexplored. Unlike …

Video Generation

VideoWeave: A Data-Centric Approach for Efficient Video Understanding

2026-01-09 · Zane Durante, Silky Singh, Arpandeep Khatua, Shobhit Agarwal 외 arxiv

Training video-language models is often prohibitively expensive due to the high cost of processing long frame sequences and the limited availability of annotated long videos. We present VideoWeave, a simple yet effective…

Video Question Answering

Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video Generation

2026-07-06 · Jingyi Lu, Kai Han arxiv

Monocular-to-stereo conversion synthesizes stereoscopic content from 2D videos for immersive 3D experiences. In modern Depth-Image-Based Rendering (DIBR) approaches, stereo inpainting of disocclusions is the critical bot…

Self-Supervised LearningVideo Generation

VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents

2026-03-26 · George Eskandar, Fengyi Shen, Mohammad Altillawi, Dong Chen 외 arxiv

Recent progress in video-to-video (V2V) translation has enabled realistic resimulation of embodied AI demonstrations, a capability that allows pretrained robot policies to be transferable to new environments without addi…

GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure

2025-12-25 · Leslie Gu, Junhwa Hur, Charles Herrmann, Fangneng Zhan 외 arxiv

We introduce GeCo, a geometry-grounded metric for jointly detecting geometric deformation and occlusion-inconsistency artifacts in static scenes. By fusing residual motion and depth priors, GeCo produces interpretable, d…

Video Generation