paper-with-me

홈 › Papers

Calibrating Sequence likelihood Improves Conditional Language Generation

2022-09-30 · Yao Zhao, Misha Khalman, Rishabh Joshi, Shashi Narayan, Mohammad Saleh, Peter J. Liu

Conditional language models are predominantly trained with maximum likelihood estimation (MLE), giving probability mass to sparsely observed target sequences. While MLE trained models assign high probability to plausible sequences given the context, the model probabilities often do not accurately rank-order generated sequences by quality. This has been empirically observed in beam search decoding as output quality degrading with large beam sizes, and decoding strategies benefiting from heuristics such as length normalization and repetition-blocking. In this work, we introduce sequence likelihood calibration (SLiC) where the likelihood of model generated sequences are calibrated to better align with reference sequences in the model's latent space. With SLiC, decoding heuristics become unnecessary and decoding candidates' quality significantly improves regardless of the decoding method. Furthermore, SLiC shows no sign of diminishing returns with model scale, and presents alternative ways to improve quality with limited training and inference budgets. With SLiC, we exceed or match SOTA results on a wide range of generation tasks spanning abstractive summarization, question generation, abstractive question answering and data-to-text generation, even with modest-sized models.

📄 PDF Abstract BibTeX arXiv:2210.00045

Code (0)

등록된 구현이 없습니다.

Tasks

abstractive question answeringAbstractive Text SummarizationBlockingData-to-Text GenerationQuestion AnsweringQuestion GenerationQuestion-GenerationText GenerationText Summarization

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Calibrating Likelihoods towards Consistency in Summarization Models

2023-10-12 · Polina Zablotskaia, Misha Khalman, Rishabh Joshi, Livio Baldini Soares 외

Despite the recent advances in abstractive text summarization, current summarization models still suffer from generating factually inconsistent summaries, reducing their utility for real-world application. We argue that …

Abstractive Text SummarizationNatural Language InferenceText Summarization

Adversarially Contrastive Estimation of Conditional Neural Processes

2023-03-23 · Zesheng Ye, Jing Du, Lina Yao

Conditional Neural Processes~(CNPs) formulate distributions over functions and generate function observations with exact conditional likelihoods. CNPs, however, have limited expressivity for high-dimensional observations…

Self-Calibrating Conformal Prediction

2024-02-11 · Lars van der Laan, Ahmed M. Alaa

In machine learning, model calibration and predictive inference are essential for producing reliable predictions and quantifying uncertainty to support decision-making. Recognizing the complementary roles of point and in…

Binary ClassificationConformal PredictionDecision MakingManagement+3

Calibrating Translation Decoding with Quality Estimation on LLMs

2025-04-26 · Di wu, Yibin Lei, Christof Monz

Neural machine translation (NMT) systems typically employ maximum a posteriori (MAP) decoding to select the highest-scoring translation from the distribution mass. However, recent evidence highlights the inadequacy of MA…

2kMachine TranslationNMTTranslation

Conditional Attribute Estimation with Autoregressive Sequence Models

2026-05-13 · Erica Stutz, Giacomo Marino, Daniella Meeker, Qiao Liu 외 arxiv

Generative models are often trained with a next-token prediction objective, yet many downstream applications require the ability to estimate or control sequence-level properties. Next-token prediction can lead to overfit…