paper-with-me

홈 › Papers

STaR: Scalable Task-Conditioned Retrieval for Long-Horizon Multimodal Robot Memory

2026-02-09 · Mingfeng Yuan, Hao Zhang, Mahan Mohammadi, Runhao Li, Jinjun Shan, Steven L. Waslander arxiv

Mobile robots are often deployed over long durations in diverse open, dynamic scenes, including indoor setting such as warehouses and manufacturing facilities, and outdoor settings such as agricultural and roadway operations. A core challenge is to build a scalable long-horizon memory that supports an agentic workflow for planning, retrieval, and reasoning over open-ended instructions at variable granularity, while producing precise, actionable answers for navigation. We present STaR, an agentic reasoning framework that (i) constructs a task-agnostic, multimodal long-term memory that generalizes to unseen queries while preserving fine-grained environmental semantics (object attributes, spatial relations, and dynamic events), and (ii) introduces a Scalable Task Conditioned Retrieval algorithm based on the Information Bottleneck principle to extract from long-term memory a compact, non-redundant, information-rich set of candidate memories for contextual reasoning. We evaluate STaR on NaVQA (mixed indoor/outdoor campus scenes) and WH-VQA, a customized warehouse benchmark with many visually similar objects built with Isaac Sim, emphasizing contextual reasoning. Across the two datasets, STaR consistently outperforms strong baselines, achieving higher success rates and markedly lower spatial error. We further deploy STaR on a real Husky wheeled robot in both indoor and outdoor environments, demonstrating robust long horizon reasoning, scalability, and practical utility. Project Website: https://trailab.github.io/STaR-website/

📄 PDF Abstract BibTeX arXiv:2602.09255

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models

2024-03-18 · Mingyang Song, Mao Zheng, Xuan Luo

Despite recent efforts to develop large language models with robust long-context capabilities, the lack of long-context benchmarks means that relatively little is known about their performance. To alleviate this gap, in …

4kPositionRetrieval

StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation

2025-01-10 · CVPR 2025 1 · Shangjin Zhai, Zhichao Ye, Jialin Liu, Weijian Xie 외

Recent advances in large reconstruction and generative models have significantly improved scene reconstruction and novel view generation. However, due to compute limitations, each inference with these large models is con…

Perpetual View GenerationScene Generation

RadIR: A Scalable Framework for Multi-Grained Medical Image Retrieval via Radiology Report Mining

2025-03-06 · Tengfei Zhang, Ziheng Zhao, Chaoyi Wu, Xiao Zhou 외

Developing advanced medical imaging retrieval systems is challenging due to the varying definitions of `similar images' across different medical contexts. This challenge is compounded by the lack of large-scale, high-qua…

Image RetrievalMedical Image RetrievalRetrieval

Spatio-Temporal and Clinical Conditioning for Fine-Grained Radiology Report Retrieval

2026-07-02 · P. Sloan, E. Simpson, M. Mirmehdi arxiv

Radiology is vital to modern healthcare, but rising imaging demand and persistent workforce shortages strain reporting capacity and clinical workflows. Automated radiology report generation has the potential to support r…

Effective Conditioned and Composed Image Retrieval Combining CLIP-Based Features

2022-01-01 · CVPR 2022 1 · Alberto Baldrati, Marco Bertini, Tiberio Uricchio, Alberto del Bimbo

Conditioned and composed image retrieval extend CBIR systems by combining a query image with an additional text that expresses the intent of the user, describing additional requests w.r.t. the visual content of the q…

Composed Image Retrieval (CoIR)Contrastive LearningImage RetrievalRetrieval