paper-with-me

Papers

RenderMem: Rendering as Spatial Memory Retrieval

2026-03-15 · JooHyun Park, HyeongYeop Kang arxiv

Embodied reasoning is inherently viewpoint-dependent: what is visible, occluded, or reachable depends critically on where the agent stands. However, existing spatial memory systems for embodied agents typically store either multi-view observations or object-centric abstractions, making it difficult to perform reasoning with explicit geometric grounding. We introduce RenderMem, a spatial memory framework that treats rendering as the interface between 3D world representations and spatial reasoning. Instead of storing fixed observations, RenderMem maintains a 3D scene representation and generates query-conditioned visual evidence by rendering the scene from viewpoints implied by the query. This enables embodied agents to reason directly about line-of-sight, visibility, and occlusion from arbitrary perspectives. RenderMem is fully compatible with existing vision-language models and requires no modification to standard architectures. Experiments in the AI2-THOR environment show consistent improvements on viewpoint-dependent visibility and occlusion queries over prior memory baselines.

📄 PDF Abstract BibTeX arXiv:2603.14669

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

GS^2: Graph-based Spatial Distribution Optimization for Compact 3D Gaussian Splatting

2026-04-02 · Xianben Yang, Tao Wang, Yuxuan Li, Yi Jin 외 arxiv

3D Gaussian Splatting (3DGS) has demonstrated breakthrough performance in novel view synthesis and real-time rendering. Nevertheless, its practicality is constrained by the high memory cost due to a huge number of Gaussi…

Novel View Synthesis

AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories

2026-02-16 · Zun Wang, Han Lin, Jaehong Yoon, Jaemin Cho 외 arxiv

Maintaining spatial world consistency over long horizons remains a central challenge for camera-controllable video generation. Existing memory-based approaches often condition generation on globally reconstructed 3D scen…

Depth EstimationVideo Generation

Latent Spatial Memory for Video World Models

2026-06-08 · Weijie Wang, Haoyu Zhao, Yifan Yang, Feng Chen 외 arxiv

Video world models that maintain 3D spatial consistency across generated frames typically rely on explicit point cloud memory constructed in RGB space. This design is both computationally expensive, requiring repeated re…

Video Generation

LODGE: Level-of-Detail Large-Scale Gaussian Splatting with Efficient Rendering

2025-05-29 · Jonas Kulhanek, Marie-Julie Rakotosaona, Fabian Manhardt, Christina Tsalicoglou 외

In this work, we present a novel level-of-detail (LOD) method for 3D Gaussian Splatting that enables real-time rendering of large-scale scenes on memory-constrained devices. Our approach introduces a hierarchical LOD rep…

3DGSGPUNeRF

SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA

2026-01-21 · Xinyi Zheng, Yunze Liu, Chi-Hao Wu, Fan Zhang 외 arxiv

We present SpatialMem, a memory-centric system for long-horizon, language-grounded retrieval and QA from egocentric video, where metric 3D serves as an interpretable indexing scaffold rather than an explicit mapping obje…