paper-with-me

Papers

RILA: Reflective and Imaginative Language Agent for Zero-Shot Semantic Audio-Visual Navigation

2024-01-01 · CVPR 2024 1 · Zeyuan Yang, Jiageng Liu, Peihao Chen, Anoop Cherian, Tim K. Marks, Jonathan Le Roux, Chuang Gan

We leverage Large Language Models (LLM) for zeroshot Semantic Audio Visual Navigation (SAVN). Existing methods utilize extensive training demonstrations for reinforcement learning yet achieve relatively low success rates and lack generalizability. The intermittent nature of auditory signals further poses additional obstacles to inferring the goal information. To address this challenge we present the Reflective and Imaginative Language Agent (RILA). By employing multi-modal models to process sensory data we instruct an LLM-based planner to actively explore the environment. During the exploration our agent adaptively evaluates and dismisses inaccurate perceptual descriptions. Additionally we introduce an auxiliary LLMbased assistant to enhance global environmental comprehension by mapping room layouts and providing strategic insights. Through comprehensive experiments and analysis we show that our method outperforms relevant baselines without training demonstrations from the environment and complementary semantic information.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Navigation

Similar Papers 제목 키워드 기반

Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation

2024-11-30 · Yiyuan Pan, Yunzhe Xu, Zhe Liu, Hesheng Wang

Humans navigate unfamiliar environments using episodic simulation and episodic memory, which facilitate a deeper understanding of the complex relationships between environments and objects. Developing an imaginative memo…

NavigateVision and Language Navigation

VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures

2025-03-16 · Yoo yeon Sung, Hannah Kim, Dan Zhang

AI practitioners increasingly use large language model (LLM) agents in compound AI systems to solve complex reasoning tasks, these agent executions often fail to meet human standards, leading to errors that compromise th…

Human Agent CollaborationLarge Language Model

What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles

2025-08-14 · Mengtao Zhou, Sifan Wu, Huan Zhang, Qi Sima 외 arxiv

We investigate the capacity of Large Language Models (LLMs) for imaginative reasoning--the proactive construction, testing, and revision of hypotheses in information-sparse environments. Existing benchmarks, often static…

Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation

2025-08-09 · Yue Hu, Junzhe Wu, Ruihan Xu, Hang Liu 외 arxiv

Semantic navigation requires an agent to navigate toward a specified target in an unseen environment. Employing an imaginative navigation strategy that predicts future scenes before taking action, can empower the agent t…

AFRILANGTUTOR: Advancing Language Tutoring and Culture Education in Low-Resource Languages with Large Language Models

2026-04-22 · Tadesse Destaw Belay, Shahriar Kabir Nahin, Israel Abebe Azime, Ocean Monjur 외 arxiv

How can language learning systems be developed for languages that lack sufficient training resources? This challenge is increasingly faced by developers across the African continent who aim to build AI systems capable of…