paper-with-me

홈 › Papers

Grounding Multimodal Large Language Models in Actions

2024-06-12 · Andrew Szot, Bogdan Mazoure, Harsh Agrawal, Devon Hjelm, Zsolt Kira, Alexander Toshev

Multimodal Large Language Models (MLLMs) have demonstrated a wide range of capabilities across many domains, including Embodied AI. In this work, we study how to best ground a MLLM into different embodiments and their associated action spaces, with the goal of leveraging the multimodal world knowledge of the MLLM. We first generalize a number of methods through a unified architecture and the lens of action space adaptors. For continuous actions, we show that a learned tokenization allows for sufficient modeling precision, yielding the best performance on downstream tasks. For discrete actions, we demonstrate that semantically aligning these actions with the native output token space of the MLLM leads to the strongest performance. We arrive at these lessons via a thorough study of seven action space adapters on five different environments, encompassing over 114 embodied tasks.

📄 PDF Abstract BibTeX arXiv:2406.07904

Code (0)

등록된 구현이 없습니다.

Tasks

World Knowledge

Similar Papers 제목 키워드 기반

Multi-Attribute Interactions Matter for 3D Visual Grounding

2024-01-01 · CVPR 2024 1 · Can Xu, Yuehui Han, Rui Xu, Le Hui 외

3D visual grounding aims to localize 3D objects described by free-form language sentences. Following the detection-then-matching paradigm existing methods mainly focus on embedding object attributes in unimodal featu…

3D visual groundingAttributeVisual Grounding

Temporal Grounding of Activities using Multimodal Large Language Models

2024-05-30 · Young Chol Song

Temporal grounding of activities, the identification of specific time intervals of actions within a larger event context, is a critical task in video understanding. Recent advancements in multimodal large language models…

Video Understanding

Multimodal Unified Attention Networks for Vision-and-Language Interactions

2019-08-12 · Zhou Yu, Yuhao Cui, Jun Yu, DaCheng Tao 외

Learning an effective attention mechanism for multimodal data is important in many vision-and-language tasks that require a synergic understanding of both the visual and textual contents. Existing state-of-the-art approa…

Question AnsweringVisual GroundingVisual Question AnsweringVisual Question Answering (VQA)

MUG: Interactive Multimodal Grounding on User Interfaces

2022-09-29 · Tao Li, Gang Li, Jingjie Zheng, Purple Wang 외

We present MUG, a novel interactive task for multimodal grounding where a user and an agent work collaboratively on an interface screen. Prior works modeled multimodal UI grounding in one round: the user gives a command …

Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision

2026-03-27 · Ling Li, Bowen Liu, Zinuo Zhan, Peng Jie 외 arxiv

Traditional Visual Grounding (VG) predominantly relies on textual descriptions to localize objects, a paradigm that inherently struggles with linguistic ambiguity and often ignores non-verbal deictic cues prevalent in re…

Referring ExpressionVisual Grounding