paper-with-me

홈 › Papers

TINA: Think, Interaction, and Action Framework for Zero-Shot Vision Language Navigation

2024-03-13 · Dingbang Li, Wenzhou Chen, Xin Lin

Zero-shot navigation is a critical challenge in Vision-Language Navigation (VLN) tasks, where the ability to adapt to unfamiliar instructions and to act in unknown environments is essential. Existing supervised learning-based models, trained using annotated data through reinforcement learning, exhibit limitations in generalization capabilities. Large Language Models (LLMs), with their extensive knowledge and emergent reasoning abilities, present a potential pathway for achieving zero-shot navigation. This paper presents a VLN agent based on LLMs, exploring approaches to the zero-shot navigation problem. To compensate for the shortcomings of LLMs in environmental perception, we propose the Thinking, Interacting, and Action (TINA) framework. TINA enables the agent to scrutinize perceptual information and autonomously query key clues within the environment through an introduced question-answering module, thereby aligning instructions with specific perceptual data. The navigation agent's perceptual abilities are enhanced through the TINA framework, while the explicit thought and query processes also improve the navigational procedure's explainability and transparency. We evaluate the performance of our method on the Room-to-Room dataset. The experiment results indicate that our approach improves the navigation performance of LLM-based agents. Our approach also outperformed some supervised learning-based methods, highlighting its efficacy in zero-shot navigation.

📄 PDF Abstract BibTeX arXiv:2403.08833

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVision-Language Navigation

Similar Papers 제목 키워드 기반

Rethinking the Extraction and Interaction of Multi-Scale Features for Vessel Segmentation

2020-10-09 · Yicheng Wu, Chengwei Pan, Shuqi Wang, Ming Zhang 외

Analyzing the morphological attributes of blood vessels plays a critical role in the computer-aided diagnosis of many cardiovascular and ophthalmologic diseases. Although being extensively studied, segmentation of blood …

Decoder

Thinking agents for zero-shot generalization to qualitatively novel tasks

2025-03-25 · Thomas Miconi, Kevin McKee, Yicong Zheng, Jed McCaleb

Intelligent organisms can solve truly novel problems which they have never encountered before, either in their lifetime or their evolution. An important component of this capacity is the ability to ``think'', that is, to…

Zero-shot Generalization

A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning

2025-12-16 · Zixin Zhang, Kanghao Chen, Hanqing Wang, Hongfei Zhang 외 arxiv

Affordance prediction, which identifies interaction regions on objects based on language instructions, is critical for embodied AI. Prevailing end-to-end models couple high-level reasoning and low-level grounding into a …

NavThinker: Action-Conditioned World Models for Coupled Prediction and Planning in Social Navigation

2026-03-16 · Tianshuai Hu, Zeying Gong, Lingdong Kong, XiaoDong Mei 외 arxiv

Social navigation requires robots to act safely in dynamic human environments. Effective behavior demands thinking ahead: reasoning about how the scene and pedestrians evolve under different robot actions rather than rea…

Reinforcement Learning

Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation

2025-11-20 · Ziyu Guo, Renrui Zhang, Hongyu Li, Manyuan Zhang 외 arxiv

Recent advances in visual generation have increasingly explored the integration of reasoning capabilities. They incorporate textual reasoning, i.e., think, either before (as pre-planning) or after (as post-refinement) th…

Reinforcement Learning