paper-with-me

Papers

PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning

2023-05-31 · Faeze Brahman, Chandra Bhagavatula, Valentina Pyatkin, Jena D. Hwang, Xiang Lorraine Li, Hirona J. Arai, Soumya Sanyal, Keisuke Sakaguchi, Xiang Ren, Yejin Choi

Procedural planning, which entails decomposing a high-level goal into a sequence of temporally ordered steps, is an important yet intricate task for machines. It involves integrating common-sense knowledge to reason about complex and often contextualized situations, e.g. ``scheduling a doctor's appointment without a phone''. While current approaches show encouraging results using large language models (LLMs), they are hindered by drawbacks such as costly API calls and reproducibility issues. In this paper, we advocate planning using smaller language models. We present PlaSma, a novel two-pronged approach to endow small language models with procedural knowledge and (constrained) language planning capabilities. More concretely, we develop symbolic procedural knowledge distillation to enhance the commonsense knowledge in small language models and an inference-time algorithm to facilitate more structured and accurate reasoning. In addition, we introduce a new related task, Replanning, that requires a revision of a plan to cope with a constrained situation. In both the planning and replanning settings, we show that orders-of-magnitude smaller models (770M-11B parameters) can compete and often surpass their larger teacher models' capabilities. Finally, we showcase successful application of PlaSma in an embodied environment, VirtualHome.

📄 PDF Abstract BibTeX arXiv:2305.19472

Code (1)

allenai/plasma 공식 구현 pytorch

Tasks

Common Sense ReasoningcounterfactualCounterfactual PlanningKnowledge DistillationScheduling

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Tracking Blobs in the Turbulent Edge Plasma of a Tokamak Fusion Device

2021-11-16 · Woonghee Han, Randall A. Pietersen, Rafael Villamor-Lora, Matthew Beveridge 외

The analysis of turbulence in plasmas is fundamental in fusion research. Despite extensive progress in theoretical modeling in the past 15 years, we still lack a complete and consistent understanding of turbulence in mag…

Hybridizing Physics and Neural ODEs for Predicting Plasma Inductance Dynamics in Tokamak Fusion Reactors

2023-10-30 · Allen M. Wang, Darren T. Garnier, Cristina Rea

While fusion reactors known as tokamaks hold promise as a firm energy source, advances in plasma control, and handling of events where control of plasmas is lost, are needed for them to be economical. A significant bottl…

Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning

2025-07-11 · Chan Young Park, Jillian Fisher, Marius Memmel, Dipika Khullar 외 arxiv

Large language models (LLMs) have shown promise in robotic procedural planning, yet their human-centric reasoning often omits the low-level, grounded details needed for robotic execution. Vision-language models (VLMs) of…

Machine Learning to Predict the Antimicrobial Activity of Cold Atmospheric Plasma-Activated Liquids

2022-07-25 · Mehmet Akif Ozdemir, Gizem Dilara Ozdemir, Merve Gul, Onan Guren 외

Plasma is defined as the fourth state of matter and non-thermal plasma can be produced at atmospheric pressure under a high electrical field. The strong and broad-spectrum antimicrobial effect of plasma-activated liquids…

ArticlesBIG-bench Machine Learningregression

Learning Plasma Dynamics and Robust Rampdown Trajectories with Predict-First Experiments at TCV

2025-02-17 · Allen M. Wang, Alessandro Pau, Cristina Rea, Oswin So 외

The rampdown in tokamak operations is a difficult to simulate phase during which the plasma is often pushed towards multiple instability limits. To address this challenge, and reduce the risk of disrupting operations, we…

Reinforcement Learning (RL)