paper-with-me

Papers Scene Generation

“Scene Generation” 태그가 달린 논문 524편 · 필터 해제

StreetDiff: Multi-view Street Scenes Generation via Cross-view Consistent Multi-view Stable Diffusion with Structure Prompts

2026-09-09 · Qi Zhang, Yanyifan Wang, Weiyuan Zhang, Hui Huang arxiv

Multi-view diffusion models have shown strong performance in scenes with strong geometric priors and sparse semantics, such as indoor rooms or simple outdoor environments (e.g., fields, courtyards). However, they often f…

Scene Generation

SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

2026-09-04 · Xingjian Ran, Xiaoye Mo, Sihao Liu, Jianyu Zhang 외 hf

Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet their scalable generation remains challenging. Recent agentic text-to-3D scene pipelines that rely on vision-lan…

Scene Generation

ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation

2026-08-31 · Jiawei Zhang, Hongsong Wang, Pan Zhou arxiv

Text-driven 3D indoor scene generation has advanced from dataset-bound layout modeling to open-vocabulary synthesis with large language and vision-language models. Yet existing methods remain limited: one-pass generators…

Scene Generation

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

2026-08-27 · Chuan Fang, Lingteng Qiu, Yixun Liang, Rui Chen 외 arxiv

Explorable image-to-scene generation is essential for applications in gaming, robotics, and virtual reality. Existing methods based on video diffusion model (VDM) commonly rely on incomplete conditioning signals such as …

Scene GenerationPoint Clouds

Towards Surgical World-Action Modeling: A Preliminary Joint Visual-Trajectory Forecasting for Surgical Motion Planning

2026-08-20 · Weiliang Huang, Huanrong Liu, Bob Zhang, Qi Dou 외 arxiv

Reliable surgical planning requires models to anticipate not only how instruments will move, but also how the operative visual state will evolve together with such motion. Existing approaches typically treat future scene…

Trajectory ForecastingTrajectory PredictionMotion ForecastingScene Generation

Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction

2026-08-19 · Zijian Xiao, Zipeng Ye, Jinkun Hao, Xiong Yang 외 arxiv

Indoor scene synthesis provides essential environments for embodied AI, robotic manipulation, and simulation-based policy learning. Recent code-based scene generation methods produce editable and extensible environments,…

Indoor Scene SynthesisScene Generation

GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

2026-08-18 · Ming Qian, Zijian Wang, Minchao Sun, Jincheng Xiong 외 arxiv

Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) reconstructions are unordered, spatially irregular, and vary widely in primitive count. We present GS-V…

Scene Generation

Training with synthetic data for drone detection in thermal imagery

2026-08-18 · Tanel Liiv, Sander Soodla, Nzamba Bignoumba, Alma M. Liezenga 외 arxiv

Ground-to-Air (G2A) drone detection in medium- and long-wave infrared (MWIR/LWIR) imagery is challenging due to reduced texture information, sensor noise, weak thermal contrast, and the scarcity of annotated data. This w…

Scene Generation

GaussianDWM++: Language-Grounded 3D Gaussian Driving World Model for Unified Scene Understanding, Editing, and Multi-Modal Generation

2026-08-17 · Tianchen Deng, Xuefeng Chen, Shuang Wu, Qu Chen 외 arxiv

Driving World Models (DWMs) have recently advanced rapidly with generative models, yet most existing methods mainly focus on conditional scene generation and lack explicit 3D scene understanding, language-grounded reason…

Scene UnderstandingScene GenerationVisual Grounding

iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

2026-08-06 · Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel, Ajad Chhatkuli 외 arxiv

Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional cons…

Synthetic Data GenerationReinforcement LearningScene Generation

WorldClaw: Agentic 3D Open-World Generation at Scale

2026-08-05 · Chunchao Guo, Jinpeng Li, Yang Li, Zilong Huang hf

Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstrea…

Scene Generation

Long-term Traffic Scene Prediction via Polynomial Representations in Autonomous Driving

2026-08-04 · Yue Yao arxiv

This thesis addresses fundamental challenges in traffic scene prediction for autonomous driving by introducing robust and computationally efficient models based on polynomial representations. While conventional sequence-…

Computational EfficiencyAutonomous DrivingScene Generation

Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation

2026-08-04 · Jialu Huang, Yingxuan You, Fei Wang, Zheng Dang arxiv

We study open-vocabulary 3D indoor layout generation, which synthesizes diverse and physically plausible scenes from unlabeled 3D assets using free-form language instructions. Recent methods leverage large language model…

Scene Generation

BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

2026-07-29 · Zhe Liu, Quan Lu, Zhaohui Du, Zhe Wang 외 arxiv

Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center o…

Scene Generation

FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows

2026-07-29 · Wenzhe He, Meng Wang, JiaWei Qian, Jinfeng Xu 외 arxiv

Existing point-based generative methods for outdoor scenes primarily focus on LiDAR-conditioned completion. During training, noisy point clouds are constructed by perturbing complete ground-truth scenes, whereas during i…

Scene GenerationPoint Clouds

PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation

2026-07-26 · Pengyu Zeng, Yuqin Dai, Jun Yin, Ziyang Han 외 arxiv

Two structural insights have been overlooked in automated residential floor plan generation. First, design is inherently progressive. Architects begin with rough strokes and refine them over time, whereas existing method…

Spatial ReasoningScene Generation

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World

2026-07-19 · Qing Zong, Yue Guo, Mengxin Yang, Yiwen Guo 외 hf

This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as static persona imitatio…

Scene Generation

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

2026-07-13 · Xinghang Li, Jun Guo, Qiwei Li, Long Qian 외 hf

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coh…

Text-to-Image GenerationScene GenerationVideo GenerationImage Editing

RoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and Evaluation

2026-07-07 · Shujie Zhang, Jingkun Yi, Weipeng Zhong, Zirui Zhou 외 arxiv

Recovering real-world scenes as interactive simulation environments can enable generalizable robot learning and reproducible policy evaluation. However, constructing scenes that are both physically stable and visually fa…

Synthetic Data GenerationScene Generation

SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

2026-07-06 · Paul Engstler, Iro Laina, Christian Rupprecht, Andrea Vedaldi arxiv

We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained layout control. Building on the ability of current image-to-3D generators to produce complex 3D assets …

Scene Generation
1–20 / 524 다음 →