paper-with-me

Papers

Steerable Scene Generation with Post Training and Inference-Time Search

2025-05-07 · Nicholas Pfaff, Hongkai Dai, Sergey Zakharov, Shun Iwase, Russ Tedrake

Training robots in simulation requires diverse 3D scenes that reflect the specific challenges of downstream tasks. However, scenes that satisfy strict task requirements, such as high-clutter environments with plausible spatial arrangement, are rare and costly to curate manually. Instead, we generate large-scale scene data using procedural models that approximate realistic environments for robotic manipulation, and adapt it to task-specific goals. We do this by training a unified diffusion-based generative model that predicts which objects to place from a fixed asset library, along with their SE(3) poses. This model serves as a flexible scene prior that can be adapted using reinforcement learning-based post training, conditional generation, or inference-time search, steering generation toward downstream objectives even when they differ from the original data distribution. Our method enables goal-directed scene synthesis that respects physical feasibility and scales across scene types. We introduce a novel MCTS-based inference-time search strategy for diffusion models, enforce feasibility via projection and simulation, and release a dataset of over 44 million SE(3) scenes spanning five diverse environments. Website with videos, code, data, and model weights: https://steerable-scene-generation.github.io/

📄 PDF Abstract BibTeX arXiv:2505.04831

Code (1)

nepfaff/steerable-scene-generation 공식 구현 pytorch

Tasks

Scene Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Urban Architect: Steerable 3D Urban Scene Generation with Layout Prior

2024-04-10 · Fan Lu, Kwan-Yee Lin, Yan Xu, Hongsheng Li 외

Text-to-3D generation has achieved remarkable success via large-scale text-to-image diffusion models. Nevertheless, there is no paradigm for scaling up the methodology to urban scale. Urban scenes, characterized by numer…

3D GenerationModel OptimizationScene GenerationText to 3D

Bayesian 3D Steerable CNNs: Enabling Equivariance and Uncertainty Quantification Simultaneously

2026-06-13 · Abhishek Keripale, Ponkrshnan Thiagarajan, Susanta Ghosh arxiv

Steerable convolutional neural networks (Steerable-CNNs) guarantee SE(3)-equivariance by parameterizing kernels as linear combinations of steerable basis functions, but their deterministic nature precludes uncertainty qu…

Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning

2024-07-22 · Kaiwen Wang, Rahul Kidambi, Ryan Sullivan, Alekh Agarwal 외

Reward-based finetuning is crucial for aligning language policies with intended behaviors (e.g., creativity and safety). A key challenge is to develop steerable language models that trade-off multiple (conflicting) objec…

Uirapuru: Timely Video Analytics for High-Resolution Steerable Cameras on Edge Devices

2025-09-01 · Guilherme H. Apostolo, Pablo Bauszat, Vinod Nigade, Henri E. Bal 외 arxiv

Real-time video analytics on high-resolution cameras has become a popular technology for various intelligent services like traffic control and crowd monitoring. While extensive work has been done on improving analytics a…

Fully Steerable 3D Spherical Neurons

2021-09-29 · Pavlo Melnyk, Michael Felsberg, Mårten Wadenbäck

Emerging from low-level vision theory, steerable filters found their counterpart in prior work on steerable convolutional neural networks equivariant to rigid transformations. In our work, we propose a steerable feed-for…

3D geometry