paper-with-me

홈 › Papers

A Picture is Worth a Thousand Words: Language Models Plan from Pixels

2023-03-16 · Anthony Z. Liu, Lajanugen Logeswaran, Sungryull Sohn, Honglak Lee

Planning is an important capability of artificial agents that perform long-horizon tasks in real-world environments. In this work, we explore the use of pre-trained language models (PLMs) to reason about plan sequences from text instructions in embodied visual environments. Prior PLM based approaches for planning either assume observations are available in the form of text (e.g., provided by a captioning model), reason about plans from the instruction alone, or incorporate information about the visual environment in limited ways (such as a pre-trained affordance function). In contrast, we show that PLMs can accurately plan even when observations are directly encoded as input prompts for the PLM. We show that this simple approach outperforms prior approaches in experiments on the ALFWorld and VirtualHome benchmarks.

📄 PDF Abstract BibTeX arXiv:2303.09031

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

One Picture is Worth a Thousand Words: A New Wallet Recovery Process

2022-05-05 · Hervé Chabannne, Vincent Despiegel, Linda Guiga

We introduce a new wallet recovery process. Our solution associates 1) visual passwords: a photograph of a secretly picked object (Chabanne et al., 2013) with 2) ImageNet classifiers transforming images into binary vecto…

Retrieval

Image and Information

2016-02-03 · Frank Nielsen

A well-known old adage says that {\em "A picture is worth a thousand words!"} (attributed to the Chinese philosopher Confucius ca 500 years BC). But more precisely, what do we mean by information in images? And how can i…

ClipMatrix: Text-controlled Creation of 3D Textured Meshes

2021-09-27 · Nikolay Jetchev

If a picture is worth thousand words, a moving 3d shape must be worth a million. We build upon the success of recent generative methods that create images fitting the semantics of a text prompt, and extend it to the cont…

MOTIF: Contextualized Images for Complex Words to Improve Human Reading

2022-06-01 · LREC 2022 6 · Xintong Wang, Florian Schneider, Özge Alacam, Prateek Chaudhury 외

MOTIF (MultimOdal ConTextualized Images For Language Learners) is a multimodal dataset that consists of 1125 comprehension texts retrieved from Wikipedia Simple Corpus. Allowing multimodal processing or enriching the con…

Reading Comprehension

Universal Differential Equations for Scientific Machine Learning

2020-01-13 · Christopher Rackauckas, Yingbo Ma, Julius Martensen, Collin Warner 외

In the context of science, the well-known adage "a picture is worth a thousand words" might well be "a model is worth a thousand datasets." In this manuscript we introduce the SciML software ecosystem as a tool for mixin…

BIG-bench Machine LearningGPU