paper-with-me

홈 › Papers

SELU: Self-Learning Embodied MLLMs in Unknown Environments

2024-10-04 · Boyu Li, Haobin Jiang, Ziluo Ding, Xinrun Xu, Haoran Li, Dongbin Zhao, Zongqing Lu

Recently, multimodal large language models (MLLMs) have demonstrated strong visual understanding and decision-making capabilities, enabling the exploration of autonomously improving MLLMs in unknown environments. However, external feedback like human or environmental feedback is not always available. To address this challenge, existing methods primarily focus on enhancing the decision-making capabilities of MLLMs through voting and scoring mechanisms, while little effort has been paid to improving the environmental comprehension of MLLMs in unknown environments. To fully unleash the self-learning potential of MLLMs, we propose a novel actor-critic self-learning paradigm, dubbed SELU, inspired by the actor-critic paradigm in reinforcement learning. The critic employs self-asking and hindsight relabeling to extract knowledge from interaction trajectories collected by the actor, thereby augmenting its environmental comprehension. Simultaneously, the actor is improved by the self-feedback provided by the critic, enhancing its decision-making. We evaluate our method in the AI2-THOR and VirtualHome environments, and SELU achieves critic improvements of approximately 28% and 30%, and actor improvements of about 20% and 24% via self-learning.

📄 PDF Abstract BibTeX arXiv:2410.03303

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingSelf-Learning

Methods 이 논문이 사용한 방법론

22 Ways to Contact: How Can I Speak to Someone at Expedia 21 Ways to Contact: How Can I Speak to Someone at Expedia, call +1-805-330-4056 or use the app’s live chat. Visit Expedia.com/contact +1-805-330-4056 to log in and request a…
Focus 설명 없음
Self-Learning 설명 없음

Similar Papers 제목 키워드 기반

Redefining Self-Normalization Property

2021-01-01 · Zhaodong Chen, Zhao WeiQin, Lei Deng, Guoqi Li 외

The approaches that prevent gradient explosion and vanishing have boosted the performance of deep neural networks in recent years. A unique one among them is the self-normalizing neural network (SNN), which is generally …

Data Augmentation

MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror

2026-04-16 · Shengyu Guo, Tongrui Ye, Jianbo Zhang, Zicheng Zhang 외 arxiv

Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated remarkable advances in perception and reasoning, suggesting their potential for embodied intelligence. While recent studies have evaluated embo…

ESCA: Contextualizing Embodied Agents via Scene-Graph Generation

2025-10-11 · Jiani Huang, Amish Sethi, Matthew Kuo, Mayank Keoliya 외 arxiv

Multi-modal large language models (MLLMs) are making rapid progress toward general-purpose embodied agents. However, existing MLLMs do not reliably capture fine-grained links between low-level visual features and high-le…

Scene Graph Generation

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

2026-07-14 · Zhishan Zou, Guoyan Sun, Zhiwei Wei, Jiancheng Pan 외 arxiv

Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also ma…

MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments

2026-06-30 · Qingyun Liu, Jiwen Zhang, Jingyi Hu, Siyuan Wang 외 arxiv

Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded environments remains underexplored. To address this gap, we introduce MECoBe…