paper-with-me

Papers

Self-Evolving 3D Scene Generation from a Single Image

2025-12-09 · Kaizhi Zheng, Yue Fan, Jing Gu, Zishuo Xu, Xuehai He, Xin Eric Wang arxiv

Generating high-quality, textured 3D scenes from a single image remains a fundamental challenge in vision and graphics. Recent image-to-3D generators recover reasonable geometry from single views, but their object-centric training limits generalization to complex, large-scale scenes with faithful structure and texture. We present EvoScene, a self-evolving, training-free framework that progressively reconstructs complete 3D scenes from single images. The key idea is combining the complementary strengths of existing models: geometric reasoning from 3D generation models and visual knowledge from video generation models. Through three iterative stages--Spatial Prior Initialization, Visual-guided 3D Scene Mesh Generation, and Spatial-guided Novel View Generation--EvoScene alternates between 2D and 3D domains, gradually improving both structure and appearance. Experiments on diverse scenes demonstrate that EvoScene achieves superior geometric stability, view-consistent textures, and unseen-region completion compared to strong baselines, producing ready-to-use 3D meshes for practical applications.

📄 PDF Abstract BibTeX arXiv:2512.08905

Code (0)

등록된 구현이 없습니다.

Tasks

Scene GenerationVideo Generation3D Generation

Similar Papers 제목 키워드 기반

Scene2Demo: Self-Evolving Embodied Data Generation via Object-Action Graph

2026-02-12 · Xiang Liu, Sen Cui, Guocai Yao, Zhong Cao 외 arxiv

We present Scene2Demo, a self-evolving framework for offline embodied data generation. Given a single real-world RGB image and a user query, Scene2Demo constructs an interactive simulated scene and generates executable t…

DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation

2026-06-30 · Siyu Yan, Yizhen Gao, Yilin Wang, Dongxing Mao 외 hf

Text-rich image generation is one of the most challenging settings in image generation, since models must simultaneously produce visually realistic images and render legible, semantically aligned, and layout-consistent t…

Image Generation

EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory

2025-10-01 · Jiahao Wang, Luoxin Ye, TaiMing Lu, Junfei Xiao 외 arxiv

Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world model that bridges panoramic video genera…

3D ReconstructionVideo Generation

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

2026-06-25 · Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar, Rao Muhammad Anwer 외 arxiv

Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and self-consistency reward schemes in existin…

Visual Question AnsweringImage CaptioningVisual Reasoning

Autonomous Video Generation with Counterfactual Controllability for Self-Evolving World Models

2026-06-23 · Xin Wang, Wenxuan Liu, Tongtong Feng, Wenwu Zhu arxiv

Existing literature claims that video generation essentially is world modelling. On the one hand, the claim is productive because it pushes generative AI beyond static images and toward temporally extended physical scene…

Video Generation