paper-with-me

Papers

Learning a Visually Grounded Memory Assistant

2022-10-07 · Meera Hahn, Kevin Carlberg, Ruta Desai, James Hillis

We introduce a novel interface for large scale collection of human memory and assistance. Using the 3D Matterport simulator we create a realistic indoor environments in which we have people perform specific embodied memory tasks that mimic household daily activities. This interface was then deployed on Amazon Mechanical Turk allowing us to test and record human memory, navigation and needs for assistance at a large scale that was previously impossible. Using the interface we collect the `The Visually Grounded Memory Assistant Dataset' which is aimed at developing our understanding of (1) the information people encode during navigation of 3D environments and (2) conditions under which people ask for memory assistance. Additionally we experiment with with predicting when people will ask for assistance using models trained on hand-selected visual and semantic features. This provides an opportunity to build stronger ties between the machine-learning and cognitive-science communities through learned models of human perception, memory, and cognition.

📄 PDF Abstract BibTeX arXiv:2210.03787

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

A Grounded Memory System For Smart Personal Assistants

2025-05-09 · Felix Ocker, Jörg Deigmöller, Pavel Smirnov, Julian Eggert

A wide variety of agentic AI applications - ranging from cognitive assistants for dementia patients to robotics - demand a robust memory system grounded in reality. In this paper, we propose such a memory system consisti…

Entity DisambiguationImage CaptioningQuestion AnsweringRetrieval+1

Said Aloud, Read Different: Cross-Modal Instability in Multimodal Models

2026-08-27 · Basel Mousi, Fahim Dalvi, Shammur Chowdhury, Firoj Alam 외 arxiv

Multimodal foundation models are increasingly used in speech-first assistants that must interpret spoken queries and produce visually grounded decisions. Yet it remains unclear whether semantically equivalent queries yie…

SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory

2026-05-30 · Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx, James Fort 외 arxiv

AI glasses present a compelling platform for AI agents to serve as personalized memory assistants. To be genuinely useful, such systems must move beyond short-term video comprehension and address memory gaps that humans …

Visual Question AnsweringAction Recognition

MobileMem: Learning from a Year of Mobile Experiences

2026-08-11 · Xinle Deng, Yida Xue, Xiangyuan Ru, Haoming Xu 외 hf

The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persistent personal assistants that can understand, remember, and continuously learn from users' experiences. S…

Information Retrieval

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents

2026-05-18 · Pan Wang, Yihao Hu, Xiujin Liu, Jingchu Yang 외 arxiv

Vision-language model (VLM) agents increasingly rely on memory-augmented reinforcement learning to reuse experience across long-horizon tasks, yet most existing frameworks store memory as text and depend on proprietary t…

Reinforcement LearningDecision Making