paper-with-me

홈 › Papers

What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory

2026-06-09 · Doeon Kwon, Junho Bang arxiv

Language-agent "memory palace" systems anchor each memory to a world coordinate, on the intuition that geometry adds something text cannot. We make that intuition testable and report three results. First, the memory-palace default of folding spatial proximity into a linear blend beside recency and importance does not help and can hurt: in a pre-registered recall experiment the shipped blend fails its own frozen test (mean Delta-Hit@5 -0.0375, Wilcoxon p=0.306), sitting at a position-blind baseline, while a geometry-led weighting wins decisively (+0.3208, p<10^-15): geometry must lead recall when the query regime is spatial. Second, memory recall and visibility must be separated: recall is occlusion-blind by design (you correctly remember the next room behind a wall), while visibility is a perception predicate over stored geometry that the live system never computed. A one-line ray-versus-voxel digital differential analyzer (DDA), re-pointed from the gaze ray the agent already casts, supplies it: text and the live FoV cone both score 0.000 on 849 behind-wall targets while cone-plus-DDA reaches 0.982 (exact McNemar p<10^-6); coordinate recall separately resolves near-duplicate locations a cosine null cannot (1.000 vs 0.533, n=150). Third, the visibility predicate is confirmed live under a git-committed pre-registration (SPMEM-OCC-LIVE-v1: eight scripted worlds, automated oracle scoring, 96 behind-wall targets, false-visible 1.000->0.000, pooled exact McNemar p=2.5x10^-29), a run that surfaced and fixed a real relay anchor defect. We concede that occlusion-needs-geometry is near-tautological; the contribution is the measurement and isolation, separating what spatial memory must store from how it is read. These pilots power a frozen confirmatory study (SPMEM-ZERO-REAL-PREREG-v1); the full human-authored multi-world study with blind raters remains future work.

📄 PDF Abstract BibTeX arXiv:2606.10299

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RenderMem: Rendering as Spatial Memory Retrieval

2026-03-15 · JooHyun Park, HyeongYeop Kang arxiv

Embodied reasoning is inherently viewpoint-dependent: what is visible, occluded, or reachable depends critically on where the agent stands. However, existing spatial memory systems for embodied agents typically store eit…

Spatial Reasoning

MUlti-Store Tracker (MUSTer): A Cognitive Psychology Inspired Approach to Object Tracking

2015-06-01 · CVPR 2015 6 · Zhibin Hong, Zhe Chen, Chaohui Wang, Xue Mei 외

Variations in the appearance of a tracked object, such as changes in geometry/photometry, camera viewpoint, illumination, or partial occlusion, pose a major challenge to object tracking. Here, we adopt cognitive psycholo…

ObjectObject Tracking

What Must Generalist Agents Remember?

2026-06-17 · Khurram Yamin, Namrata Deka, Maitreyi Swaroop, Albert Ting 외 arxiv

This paper develops a formal account of what generalist agents must store in memory in order to act near-optimally across multiple environments and goals. It shows that when two domains share an observational bottleneck …

Video-based Person Re-identification with Spatial and Temporal Memory Networks

2021-08-20 · ICCV 2021 10 · Chanho Eom, Geon Lee, Junghyup Lee, Bumsub Ham

Video-based person re-identification (reID) aims to retrieve person videos with the same identity as a query person across multiple cameras. Spatial and temporal distractors in person videos, such as background clutter a…

Person Re-IdentificationVideo-Based Person Re-Identification

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

2026-07-27 · Ruizhe Li, Mingxuan Du, Benfeng Xu, Zhendong Mao hf

Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will r…