paper-with-me

Papers

ArbiViewGen: Controllable Arbitrary Viewpoint Camera Data Generation for Autonomous Driving via Stable Diffusion Models

2025-08-07 · Yatong Lan, Jingfeng Chen, Yiru Wang, Lei He arxiv

Arbitrary viewpoint image generation holds significant potential for autonomous driving, yet remains a challenging task due to the lack of ground-truth data for extrapolated views, which hampers the training of high-fidelity generative models. In this work, we propose Arbiviewgen, a novel diffusion-based framework for the generation of controllable camera images from arbitrary points of view. To address the absence of ground-truth data in unseen views, we introduce two key components: Feature-Aware Adaptive View Stitching (FAVS) and Cross-View Consistency Self-Supervised Learning (CVC-SSL). FAVS employs a hierarchical matching strategy that first establishes coarse geometric correspondences using camera poses, then performs fine-grained alignment through improved feature matching algorithms, and identifies high-confidence matching regions via clustering analysis. Building upon this, CVC-SSL adopts a self-supervised training paradigm where the model reconstructs the original camera views from the synthesized stitched images using a diffusion model, enforcing cross-view consistency without requiring supervision from extrapolated data. Our framework requires only multi-camera images and their associated poses for training, eliminating the need for additional sensors or depth maps. To our knowledge, Arbiviewgen is the first method capable of controllable arbitrary view camera image generation in multiple vehicle configurations.

📄 PDF Abstract BibTeX arXiv:2508.05236

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningAutonomous DrivingImage Generation

Similar Papers 제목 키워드 기반

FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint

2025-12-12 · Jiapeng Tang, Kai Li, Chengxiang Yin, Liuhao Ge 외 arxiv

We introduce FactorPortrait, a video diffusion method for controllable portrait animation that enables lifelike synthesis from disentangled control signals of facial expressions, head movement, and camera viewpoints. Giv…

Novel View Synthesis

FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis

2025-02-12 · CVPR 2025 1 · Wonjoon Jin, Qi Dai, Chong Luo, Seung-Hwan Baek 외

This paper presents FloVD, a novel optical-flow-based video diffusion model for camera-controllable video generation. FloVD leverages optical flow maps to represent motions of the camera and moving objects. This approach…

Motion SynthesisOptical Flow EstimationVideo Generation

ReRoPE: Repurposing RoPE for Relative Camera Control

2026-02-08 · Chunyang Li, Yuanbo Yang, Jiahao Shao, Hongyu Zhou 외 arxiv

Video generation with controllable camera viewpoints is essential for applications such as interactive content creation, gaming, and simulation. Existing methods typically adapt pre-trained video models using camera pose…

Video Generation

3DCarGen: Scalable 3D Car Generation via 3D-consistent Multi-view Synthesis

2026-06-23 · Hongli Xiao, Youjian Zhang, Yaohui Jin, Xiaoguang Ren 외 arxiv

High-quality 3D vehicle assets are essential for autonomous driving simulation. Although multi-view diffusion-based paradigms enable controllable single-image reconstruction, they typically produce limited viewpoints and…

Image ReconstructionAutonomous Driving

SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification

2025-04-13 · Xiang Hu, Pingping Zhang, Yuhao Wang, Bin Yan 외

Aerial-Ground Person Re-IDentification (AG-ReID) aims to retrieve specific persons across cameras with different viewpoints. Previous works focus on designing discriminative ReID models to maintain identity consistency d…

Person Re-Identification