paper-with-me

홈 › Papers

Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks

2024-07-29 · Marco AF Pimentel, Clément Christophe, Tathagata Raha, Prateek Munjal, Praveen K Kanithi, Shadab Khan

As large language models (LLMs) continue to evolve, the need for robust and standardized evaluation benchmarks becomes paramount. Evaluating the performance of these models is a complex challenge that requires careful consideration of various linguistic tasks, model architectures, and benchmarking methodologies. In recent years, various frameworks have emerged as noteworthy contributions to the field, offering comprehensive evaluation tests and benchmarks for assessing the capabilities of LLMs across diverse domains. This paper provides an exploration and critical analysis of some of these evaluation methodologies, shedding light on their strengths, limitations, and impact on advancing the state-of-the-art in natural language processing.

📄 PDF Abstract BibTeX arXiv:2407.21072

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingLanguage Model EvaluationLanguage ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Exploring Variability in Fine-Tuned Models for Text Classification with DistilBERT

2024-12-31 · Giuliano Lorenzoni, Ivens Portugal, Paulo Alencar, Donald Cowan

This study evaluates fine-tuning strategies for text classification using the DistilBERT model, specifically the distilbert-base-uncased-finetuned-sst-2-english variant. Through structured experiments, we examine the inf…

regressionSST-2text-classificationText Classification

Bias, Consistency, and Partisanship in U.S. Asylum Cases: A Machine Learning Analysis of Extraneous Factors in Immigration Court Decisions

2023-05-25 · Vyoma Raman, Catherine Vera, CJ Manna

In this study, we introduce a novel two-pronged scoring system to measure individual and systemic bias in immigration courts under the U.S. Executive Office of Immigration Review (EOIR). We analyze nearly 6 million immig…

Decision MakingTime SeriesTime Series Analysis

Quantifying Fairness in LLMs Beyond Tokens: A Semantic and Statistical Perspective

2025-06-23 · Weijie Xu, Yiwen Wang, Chi Xue, Xiangkun Hu 외

Large Language Models (LLMs) often generate responses with inherent biases, undermining their reliability in real-world applications. Existing evaluation methods often overlook biases in long-form responses and the intri…

counterfactualFairness

Benchmarking atmospheric circulation variability in an AI emulator, ACE2, and a hybrid model, NeuralGCM

2025-10-06 · Ian Baxter, Hamid Pahlavan, Pedram Hassanzadeh, Katharine Rucker 외 arxiv

Physics-based atmosphere-land models with prescribed sea surface temperature have notable successes but also biases in their ability to represent atmospheric variability compared to observations. Recently, AI emulators a…

Reconsideration on evaluation of machine learning models in continuous monitoring using wearables

2023-12-04 · Cheng Ding, Zhicheng Guo, Cynthia Rudin, Ran Xiao 외

This paper explores the challenges in evaluating machine learning (ML) models for continuous health monitoring using wearable devices beyond conventional metrics. We state the complexities posed by real-world variability…