paper-with-me

홈 › Papers

JOG3R: Towards 3D-Consistent Video Generators

2025-01-02 · Chun-Hao Paul Huang, Niloy Mitra, Hyeonho Jeong, Jae Shin Yoon, Duygu Ceylan

Emergent capabilities of image generators have led to many impactful zero- or few-shot applications. Inspired by this success, we investigate whether video generators similarly exhibit 3D-awareness. Using structure-from-motion as a 3D-aware task, we test if intermediate features of a video generator - OpenSora in our case - can support camera pose estimation. Surprisingly, at first, we only find a weak correlation between the two tasks. Deeper investigation reveals that although the video generator produces plausible video frames, the frames themselves are not truly 3D-consistent. Instead, we propose to jointly train for the two tasks, using photometric generation and 3D aware errors. Specifically, we find that SoTA video generation and camera pose estimation (i.e.,DUSt3R [79]) networks share common structures, and propose an architecture that unifies the two. The proposed unified model, named \nameMethod, produces camera pose estimates with competitive quality while producing 3D-consistent videos. In summary, we propose the first unified video generator that is 3D-consistent, generates realistic video frames, and can potentially be repurposed for other 3D-aware tasks.

📄 PDF Abstract BibTeX arXiv:2501.01409

Code (0)

등록된 구현이 없습니다.

Tasks

Camera Pose EstimationPose EstimationVideo Generation

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration

2026-05-14 · Ruicheng Zhang, Kaixi Cong, Jun Zhou, Zhizhou Zhong 외 arxiv

Aligning streaming autoregressive (AR) video generators with human preferences is challenging. Existing reinforcement learning methods predominantly rely on noise-based exploration and SDE-based surrogate policies that a…

Reinforcement LearningVideo Alignment

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

2026-08-14 · Shuo Liang, Yixing Ma, Pengfei Zhou, Xingyan Chen 외 arxiv

Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited …

OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

2026-07-09 · Hongyu Liu, Chun Wang, Feng Gao, Xuanhua He 외 arxiv

We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low latency, but still …

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

2026-05-11 · Keming Wu, Yijing Cui, Wenhan Xue, Qijie Wang 외 arxiv

Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark th…

Video Generation

WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling

2025-12-08 · Shaoheng Fang, Hanwen Jiang, Yunpeng Bai, Niloy J. Mitra 외 arxiv

Recent video generators achieve striking photorealism, yet remain fundamentally inconsistent in 3D. We present WorldReel, a 4D video generator that is natively spatio-temporally consistent. WorldReel jointly produces RGB…

Video Generation