paper-with-me

홈 › Papers

Impossible Triangle: What's Next for Pre-trained Language Models?

2022-04-13 · Chenguang Zhu, Michael Zeng

Recent development of large-scale pre-trained language models (PLM) have significantly improved the capability of models in various NLP tasks, in terms of performance after task-specific fine-tuning and zero-shot / few-shot learning. However, many of such models come with a dauntingly huge size that few institutions can afford to pre-train, fine-tune or even deploy, while moderate-sized models usually lack strong generalized few-shot learning capabilities. In this paper, we first elaborate the current obstacles of using PLM models in terms of the Impossible Triangle: 1) moderate model size, 2) state-of-the-art few-shot learning capability, and 3) state-of-the-art fine-tuning capability. We argue that all existing PLM models lack one or more properties from the Impossible Triangle. To remedy these missing properties of PLMs, various techniques have been proposed, such as knowledge distillation, data augmentation and prompt learning, which inevitably brings additional work to the application of PLMs in real scenarios. We then offer insights into future research directions of PLMs to achieve the Impossible Triangle, and break down the task into several key phases.

📄 PDF Abstract BibTeX arXiv:2204.06130

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationFew-Shot LearningGeneralized Few-Shot LearningKnowledge DistillationPrompt Learning

Similar Papers 제목 키워드 기반

MeshGPT: Generating Triangle Meshes with Decoder-Only Transformers

2023-11-27 · CVPR 2024 1 · Yawar Siddiqui, Antonio Alliegro, Alexey Artemov, Tatiana Tommasi 외

We introduce MeshGPT, a new approach for generating triangle meshes that reflects the compactness typical of artist-created meshes, in contrast to dense triangle meshes extracted by iso-surfacing methods from neural fiel…

Decoder

WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models

2024-05-23 · Peng Wang, Zexi Li, Ningyu Zhang, Ziwen Xu 외

Large language models (LLMs) need knowledge updates to meet the ever-growing world facts and correct the hallucinated responses, facilitating the methods of lifelong model editing. Where the updated knowledge resides in …

HallucinationModel EditingQuestion AnsweringRetrieval

Implications of sparsity and high triangle density for graph representation learning

2022-10-27 · Hannah Sansford, Alexander Modell, Nick Whiteley, Patrick Rubin-Delanchy

Recent work has shown that sparse graphs containing many triangles cannot be reproduced using a finite-dimensional representation of the nodes, in which link probabilities are inner products. Here, we show that such grap…

Graph Representation LearningRepresentation LearningVocal Bursts Intensity Prediction

Towards the Law of Capacity Gap in Distilling Language Models

2023-11-13 · Chen Zhang, Dawei Song, Zheyu Ye, Yan Gao

Language model (LM) distillation is a trending area that aims to distil the knowledge residing in a large teacher LM to a small student one. While various methods have been proposed to maximize the effectiveness of the d…

Language Modelling

When transformers learn "impossible" languages, what do they learn?

2026-06-29 · Ram Janarthan, Coleman Haley, Sharon Goldwater arxiv

Recent work suggests that transformer language models show a bias towards human languages over unnatural ("impossible") languages argued to be unacquirable by humans. However, this literature has largely based these clai…