paper-with-me

홈 › Papers

Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data

2022-02-06 · Yaoqing Yang, Ryan Theisen, Liam Hodgkinson, Joseph E. Gonzalez, Kannan Ramchandran, Charles H. Martin, Michael W. Mahoney

Selecting suitable architecture parameters and training hyperparameters is essential for enhancing machine learning (ML) model performance. Several recent empirical studies conduct large-scale correlational analysis on neural networks (NNs) to search for effective \emph{generalization metrics} that can guide this type of model selection. Effective metrics are typically expected to correlate strongly with test performance. In this paper, we expand on prior analyses by examining generalization-metric-based model selection with the following objectives: (i) focusing on natural language processing (NLP) tasks, as prior work primarily concentrates on computer vision (CV) tasks; (ii) considering metrics that directly predict \emph{test error} instead of the \emph{generalization gap}; (iii) exploring metrics that do not need access to data to compute. From these objectives, we are able to provide the first model selection results on large pretrained Transformers from Huggingface using generalization metrics. Our analyses consider (I) hundreds of Transformers trained in different settings, in which we systematically vary the amount of data, the model size and the optimization hyperparameters, (II) a total of 51 pretrained Transformers from eight families of Huggingface NLP models, including GPT2, BERT, etc., and (III) a total of 28 existing and novel generalization metrics. Despite their niche status, we find that metrics derived from the heavy-tail (HT) perspective are particularly useful in NLP tasks, exhibiting stronger correlations than other, more popular metrics. To further examine these metrics, we extend prior formulations relying on power law (PL) spectral distributions to exponential (EXP) and exponentially-truncated power law (E-TPL) families.

📄 PDF Abstract BibTeX arXiv:2202.02842

Code (1)

nsfzyzz/generalization_metrics_for_nlp 공식 구현 pytorch

Tasks

Model Selection

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics

2021-05-19 · ACL 2021 5 · Jian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu 외

Automatic metrics are essential for developing natural language generation (NLG) models, particularly for open-ended language generation tasks such as story generation. However, existing automatic metrics are observed to…

Story GenerationText Generation

Measuring the Measuring Tools: An Automatic Evaluation of Semantic Metrics for Text Corpora

2022-11-29 · George Kour, Samuel Ackerman, Orna Raz, Eitan Farchi 외

The ability to compare the semantic similarity between text corpora is important in a variety of natural language processing applications. However, standard methods for evaluating these metrics have yet to be established…

Semantic SimilaritySemantic Textual Similarity

On the Intrinsic and Extrinsic Fairness Evaluation Metrics for Contextualized Language Representations

2022-03-25 · ACL 2022 5 · Yang Trista Cao, Yada Pruksachatkun, Kai-Wei Chang, Rahul Gupta 외

Multiple metrics have been introduced to measure fairness in various natural language processing tasks. These metrics can be roughly categorized into two categories: 1) \emph{extrinsic metrics} for evaluating fairness in…

Fairness

Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?

2025-02-07 · Sourabrata Mukherjee, Atul Kr. Ojha, John P. McCrae, Ondrej Dusek

Text Style Transfer (TST) is the task of transforming a text to reflect a particular style while preserving its original content. Evaluating TST outputs is a multidimensional challenge, requiring the assessment of style …

Machine TranslationStyle TransferText Style Transfer

DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question Answering

2023-07-13 · Pei Ke, Fei Huang, Fei Mi, Yasheng Wang 외

Existing evaluation metrics for natural language generation (NLG) tasks face the challenges on generalization ability and interpretability. Specifically, most of the well-performed metrics are required to train on evalua…

Dialogue Generationnlg evaluationQuestion AnsweringSentence+2