paper-with-me

Papers

LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions

2023-04-27 · Minghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, Alham Fikri Aji

Large language models (LLMs) with instruction fine-tuning demonstrate superior generative capabilities. However, these models are resource-intensive. To alleviate this issue, we explore distilling knowledge from instruction-tuned LLMs into much smaller ones. To this end, we carefully develop a large set of 2.58M instructions based on both existing and newly-generated instructions. In addition to being sizable, we design our instructions to cover a broad set of topics to ensure diversity. Extensive analysis of our instruction dataset confirms its diversity, and we generate responses for these instructions using gpt-3.5-turbo. Leveraging these instructions, we fine-tune a diverse herd of models, collectively referred to as LaMini-LM, which includes models from both the encoder-decoder and decoder-only families, with varying sizes. We evaluate the performance of our models using automatic metrics on 15 different natural language processing (NLP) benchmarks, as well as through human assessment. The results demonstrate that our proposed LaMini-LM models are comparable to competitive baselines, while being much smaller in size.

📄 PDF Abstract BibTeX arXiv:2304.14402

Code (1)

mbzuai-nlp/lamini-lm 공식 구현

Tasks

Common Sense ReasoningCoreference ResolutionDecoderDiversityLanguage ModellingNatural Language InferenceQuestion AnsweringSentence CompletionWord Sense Disambiguation

Similar Papers 제목 키워드 기반

Epidemic modelling of bovine tuberculosis in cattle herds and badgers in Ireland

2020-07-13

Bovine tuberculosis, a disease that affects cattle and badgers in Ireland, was studied via stochastic epidemic modeling using incidence data from the Four Area Project (Griffin et al., 2005). The Four Area Project was a …

GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks

2025-09-28 · Cong Chen, Kaixiang Ji, Hao Zhong, Muzhi Zhu 외 arxiv

Autonomous agents for long-sequence Graphical User Interface tasks are hindered by sparse rewards and the intractable credit assignment problem. To address these challenges, we introduce GUI-Shepherd, a Process Reward Mo…

A Continuification-Based Control Solution for Large-Scale Shepherding

2024-11-07 · Beniamino Di Lorenzo, Gian Carlo Maffettone, Mario di Bernardo

In this paper, we address the large-scale shepherding control problem using a continuification-based strategy. We consider a scenario in which a large group of follower agents (targets) must be confined within a designat…

Mixed Reality

Web-Shepherd: Advancing PRMs for Reinforcing Web Agents

2025-05-21 · Hyungjoo Chae, Sunghwan Kim, Junhee Cho, Seungone Kim 외

Web navigation is a unique domain that can automate many repetitive real-life tasks and is challenging as it requires long-horizon sequential decision making beyond typical multimodal large language model (MLLM) tasks. Y…

Large Language ModelMultimodal Large Language ModelSequential Decision Making

Emergent Cooperative Strategies for Multi-Agent Shepherding via Reinforcement Learning

2024-11-08 · Italo Napolitano, Andrea Lama, Francesco De Lellis, Mario di Bernardo

We present a decentralized reinforcement learning (RL) approach to address the multi-agent shepherding control problem, departing from the conventional assumption of cohesive target groups. Our two-layer control architec…

Reinforcement Learning (RL)