paper-with-me

홈 › Papers

Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs

2024-07-01 · Minh Nguyen, Andrew Baker, Clement Neo, Allen Roush, Andreas Kirsch, Ravid Shwartz-Ziv

Large Language Models (LLMs) generate text by sampling the next token from a probability distribution over the vocabulary at each decoding step. Popular sampling methods like top-p (nucleus sampling) often struggle to balance quality and diversity, especially at higher temperatures which lead to incoherent or repetitive outputs. We propose min-p sampling, a dynamic truncation method that adjusts the sampling threshold based on the model's confidence by using the top token's probability as a scaling factor. Our experiments on benchmarks including GPQA, GSM8K, and AlpacaEval Creative Writing show that min-p sampling improves both the quality and diversity of generated text across different model families (Mistral and Llama 3) and model sizes (1B to 123B parameters), especially at higher temperatures. Human evaluations further show a clear preference for min-p sampling, in both text quality and creativity. Min-p sampling has been adopted by popular open-source LLM frameworks, including Hugging Face Transformers, VLLM, and many others, highlighting its considerable impact on improving text generation quality.

📄 PDF Abstract BibTeX arXiv:2407.01082

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityGSM8KText Generation

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…
BASE 설명 없음

Similar Papers 제목 키워드 기반

ReMIND: Orchestrating Modular Large Language Models for Controllable Serendipity A REM-Inspired System Design for Emergent Creative Ideation

2026-01-12 · Makoto Sato arxiv

Large language models (LLMs) are used not only for problem solving but also for creative ideation; however, eliciting serendipitous insights that are both novel and internally coherent remains difficult. While stochastic…

THOMAS: Trajectory Heatmap Output with learned Multi-Agent Sampling

2021-10-13 · ICLR 2022 4 · Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bogdan Stanciulescu 외

In this paper, we propose THOMAS, a joint multi-agent trajectory prediction framework allowing for an efficient and consistent prediction of multi-agent multi-modal trajectories. We present a unified model architecture f…

Image GenerationPredictionTrajectory Prediction

Co-Writing Screenplays and Theatre Scripts with Language Models: An Evaluation by Industry Professionals

2022-09-29 · Piotr Mirowski, Kory W. Mathewson, Jaylen Pittman, Richard Evans

Language models are increasingly attracting interest from writers. However, such models lack long-range semantic coherence, limiting their usefulness for longform creative writing. We address this limitation by applying …

Text Generation

Creative GANs for generating poems, lyrics, and metaphors

2019-09-20 · Asir Saeed, Suzana Ilić, Eva Zangerle

Generative models for text have substantially contributed to tasks like machine translation and language modeling, using maximum likelihood optimization (MLE). However, for creative text generation, where multiple output…

Generative Adversarial NetworkLanguage ModelingLanguage ModellingMachine Translation+2

Multi-model assessment of heat decarbonisation options in the UK using electricity and hydrogen

2022-06-06 · Marko Aunedi, Maria Yliruka, Shahab Dehghan, Antonio Marco Pantaleo 외

Delivering low-carbon heat will require the substitution of natural gas with low-carbon alternatives such as electricity and hydrogen. The objective of this paper is to develop a method to soft-link two advanced, investm…