paper-with-me

홈 › Papers

SceneForge: Structured World Supervision from 3D Interventions

2026-05-14 · Jizhizi Li, Jiayang Ao, Danny Wicks, Petru-Daniel Tudosiu arxiv

Many multimodal learning tasks require supervision that remains consistent across edits, viewpoints, and scene-level interventions. However, such supervision is difficult to obtain from observation-level datasets, which do not expose the underlying scene state or how changes propagate through it. We present SceneForge, an intervention-driven framework that generates structured supervision from editable 3D world states. SceneForge represents each scene as a persistent world with semantic, geometric, and physical dependencies. By applying explicit interventions (e.g., object removal or camera variation) and propagating their effects through scene dependencies, SceneForge renders supervision that remains consistent with object structure and scene-level effects. This produces aligned outputs including counterfactual observations, multi-view observations, and effect-aware signals such as shadows and reflections, all derived from a shared world state rather than post hoc image-space processing. We instantiate SceneForge using Infinigen and Blender to construct a licensing-clean indoor supervision resource with a large number of counterfactual pairs and aligned annotations from over 2K scenes, covering both diverse single-view and registered multi-view settings. Under matched training budgets, incorporating SceneForge supervision improves both object removal and scene removal performance across multiple benchmarks in both quantitative and qualitative evaluation. These results indicate that modeling supervision as structured state transitions in editable worlds provides a practical and scalable foundation for intervention-consistent multimodal learning.

📄 PDF Abstract BibTeX arXiv:2605.14399

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SCENEFORGE: Enhancing 3D-text alignment with Structured Scene Compositions

2025-09-19 · Cristian Sbrolli, Matteo Matteucci arxiv

The whole is greater than the sum of its parts-even in 3D-text contrastive learning. We introduce SceneForge, a novel framework that enhances contrastive alignment between 3D point clouds and text through structured mult…

Visual Question AnsweringContrastive LearningSpatial ReasoningPoint Clouds

ChronoMedicalWorld: A Medical World Model for Learning Patient Trajectories from Longitudinal Care Data

2026-05-21 · Jiangyuan Wang, Xuyong Chen, Junwei He, Xu Xu 외 arxiv

Long-horizon clinical simulation -- predicting how a patient's physiology evolves over years under specified interventions -- is central to chronic-disease care, yet existing electronic health record (EHR) models are pre…

Trajectory Forecasting

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

2026-07-29 · Yiming Cai, Fangjie Yu, Meiqing Yu, Ziyue Shi 외 arxiv

World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. We introduce CG-World, a large-s…

Video Generation

ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning

2026-06-15 · Wei Xiao, Weiliang Tang, Yuying Ge, Hui Zhou 외 arxiv

Human interventions provide crucial corrective signals for post-training Vision-Language-Action (VLA) models. However, enabling seamless humanoid interventions is a formidable systems challenge due to complex whole-body …

Reinforcement Learning

Concept Embedding Models: Beyond the Accuracy-Explainability Trade-Off

2022-09-19 · Mateo Espinosa Zarlenga, Pietro Barbiero, Gabriele Ciravegna, Giuseppe Marra 외

Deploying AI-powered systems requires trustworthy models supporting effective human interactions, going beyond raw prediction accuracy. Concept bottleneck models promote trustworthiness by conditioning classification tas…