paper-with-me

홈 › Papers

DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment

2023-07-01 · Yanjiang Guo, Yen-Jen Wang, Lihan Zha, Jianyu Chen

Large language models (LLMs) encode a vast amount of semantic knowledge and possess remarkable understanding and reasoning capabilities. Previous work has explored how to ground LLMs in robotic tasks to generate feasible and executable textual plans. However, low-level execution in the physical world may deviate from the high-level textual plan due to environmental perturbations or imperfect controller design. In this paper, we propose \textbf{DoReMi}, a novel language model grounding framework that enables immediate Detection and Recovery from Misalignments between plan and execution. Specifically, we leverage LLMs to play a dual role, aiding not only in high-level planning but also generating constraints that can indicate misalignment during execution. Then vision language models (VLMs) are utilized to detect constraint violations continuously. Our pipeline can monitor the low-level execution and enable timely recovery if certain plan-execution misalignment occurs. Experiments on various complex tasks including robot arms and humanoid robots demonstrate that our method can lead to higher task success rates and shorter task completion times. Videos of DoReMi are available at \url{https://sites.google.com/view/doremi-paper}.

📄 PDF Abstract BibTeX arXiv:2307.00329

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingQuestion AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

2023-05-17 · NeurIPS 2023 11 · Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du 외

The mixture proportions of pretraining data domains (e.g., Wikipedia, books, web text) greatly affect language model (LM) performance. In this paper, we propose Domain Reweighting with Minimax Optimization (DoReMi), whic…

Language ModelingLanguage Modelling

HAR-DoReMi: Optimizing Data Mixture for Self-Supervised Human Activity Recognition Across Heterogeneous IMU Datasets

2025-03-16 · Lulu Ban, Tao Zhu, Xiangqing Lu, Qi Qiu 외

Cross-dataset Human Activity Recognition (HAR) suffers from limited model generalization, hindering its practical deployment. To address this critical challenge, inspired by the success of DoReMi in Large Language Models…

Activity RecognitionHuman Activity Recognition

DOREMI: Optimizing Long Tail Predictions in Document-Level Relation Extraction

2026-01-16 · Laura Menotti, Stefano Marchesin, Gianmaria Silvello arxiv

Document-Level Relation Extraction (DocRE) presents significant challenges due to its reliance on cross-sentence context and the long-tail distribution of relation types, where many relations have scarce training example…

Document-level Relation Extraction

DoReMi: First glance at a universal OMR dataset

2021-07-16 · Elona Shatri, György Fazekas

The main challenges of Optical Music Recognition (OMR) come from the nature of written music, its complexity and the difficulty of finding an appropriate data representation. This paper provides a first look at DoReMi, a…

object-detectionObject Detection

DoReMi: Bridging 3D Domains via Topology-Aware Domain-Representation Mixture of Experts

2025-11-14 · Mingwei Xing, Xinliang Wang, Yifeng Shi arxiv

Constructing a unified 3D scene understanding model has long been hindered by the significant topological discrepancies across different sensor modalities. While applying the Mixture-of-Experts (MoE) architecture is an e…

Scene Understanding