paper-with-me

홈 › Papers

Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning

2023-01-27 · NeurIPS 2023 11 · Xinyi Wang, Wanrong Zhu, Michael Saxon, Mark Steyvers, William Yang Wang

In recent years, pre-trained large language models (LLMs) have demonstrated remarkable efficiency in achieving an inference-time few-shot learning capability known as in-context learning. However, existing literature has highlighted the sensitivity of this capability to the selection of few-shot demonstrations. Current understandings of the underlying mechanisms by which this capability arises from regular language model pretraining objectives remain disconnected from the real-world LLMs. This study aims to examine the in-context learning phenomenon through a Bayesian lens, viewing real-world LLMs as latent variable models. On this premise, we propose an algorithm to select optimal demonstrations from a set of annotated data with a small LM, and then directly generalize the selected demonstrations to larger LMs. We demonstrate significant improvement over baselines, averaged over eight GPT models on eight real-world text classification datasets. We also demonstrate the real-world usefulness of our algorithm on GSM8K, a math word problem dataset. Our empirical findings support our hypothesis that LLMs implicitly infer a latent variable containing task information.

📄 PDF Abstract BibTeX arXiv:2301.11916

Code (1)

wangxinyilinda/concept-based-demonstration-selection 공식 구현 pytorch

Tasks

Few-Shot LearningGSM8KIn-Context LearningLanguage ModelingLanguage ModellingMathtext-classificationText Classification

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Residual Connection 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Linear Warmup With Cosine Annealing Linear Warmup With Cosine Annealing is a learning rate schedule where we increase the learning rate linearly for $n$ updates and then anneal according to a cosine schedule…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models

2024-06-21 · Mengdan Zhu, Raasikh Kanjiani, Jiahui Lu, Andrew Choi 외

Deep generative models like VAEs and diffusion models have advanced various generation tasks by leveraging latent variables to learn data distributions and generate high-quality samples. Despite the field of explainable …

Uncertainty Quantification

UniCog: Uncovering Cognitive Abilities of LLMs through Latent Mind Space Analysis

2026-01-25 · Jiayu Liu, Yinhe Long, Zhenya Huang, Enhong Chen arxiv

A growing body of research suggests that the cognitive processes of large language models (LLMs) differ fundamentally from those of humans. However, existing interpretability methods remain limited in explaining how cogn…

Explaining latent representations of generative models with large multimodal models

2024-02-02 · Mengdan Zhu, Zhenke Liu, Bo Pan, Abhinav Angirekula 외

Learning interpretable representations of data generative latent factors is an important topic for the development of artificial intelligence. With the rise of the large multimodal model, it can align images with text to…

DisentanglementExplanation Generation

Disentanglement Analysis in Deep Latent Variable Models Matching Aggregate Posterior Distributions

2025-01-26 · Surojit Saha, Sarang Joshi, Ross Whitaker

Deep latent variable models (DLVMs) are designed to learn meaningful representations in an unsupervised manner, such that the hidden explanatory factors are interpretable by independent latent variables (aka disentanglem…

Disentanglement

Structured Recognition for Generative Models with Explaining Away

2022-09-12 · Changmin Yu, Hugo Soulat, Neil Burgess, Maneesh Sahani

A key goal of unsupervised learning is to go beyond density estimation and sample generation to reveal the structure inherent within observed data. Such structure can be expressed in the pattern of interactions between e…

Density EstimationHippocampusTime Series AnalysisVariational Inference