paper-with-me

Papers

Evaluating Large Language Models with fmeval

2024-07-15 · Pola Schwöbel, Luca Franceschi, Muhammad Bilal Zafar, Keerthan Vasist, Aman Malhotra, Tomer Shenhar, Pinal Tailor, Pinar Yilmaz, Michael Diamond, Michele Donini

fmeval is an open source library to evaluate large language models (LLMs) in a range of tasks. It helps practitioners evaluate their model for task performance and along multiple responsible AI dimensions. This paper presents the library and exposes its underlying design principles: simplicity, coverage, extensibility and performance. We then present how these were implemented in the scientific and engineering choices taken when developing fmeval. A case study demonstrates a typical use case for the library: picking a suitable model for a question answering task. We close by discussing limitations and further work in the development of the library. fmeval can be found at https://github.com/aws/fmeval.

📄 PDF Abstract BibTeX arXiv:2407.12872

Code (1)

aws/fmeval 공식 구현

Tasks

Question Answering

Methods 이 논문이 사용한 방법론

Library 설명 없음

Similar Papers 제목 키워드 기반

TurkBench: A Benchmark for Evaluating Turkish Large Language Models

2026-01-11 · Çağrı Toraman, Ahmet Kaan Sever, Ayse Aysu Cengiz, Elif Ecem Arslan 외 arxiv

With the recent surge in the development of large language models, the need for comprehensive and language-specific evaluation benchmarks has become critical. While significant progress has been made in evaluating Englis…

Instruction Following

NRITYAM: Language Models Meet Art and Heritage of Dance

2026-06-18 · Punit Kumar Singh, Niladri Ghosh, Advait Joshiınst, Shailee Choudhary 외 arxiv

Language models have become essential tools in shaping modern workflows. However, their global effectiveness hinges on a nuanced understanding of local socio-cultural contexts. To address this gap, we present NRITYAM, a …

Oracle-Checker Scheme for Evaluating a Generative Large Language Model

2024-05-06 · Yueling Jenny Zeng, Li-C. Wang, Thomas Ibbetson

This work presents a novel approach called oracle-checker scheme for evaluating the answer given by a generative large language model (LLM). Two types of checkers are presented. The first type of checker follows the idea…

Language ModelingLanguage ModellingLarge Language Model

FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models

2025-04-20 · Mehrnoush Shamsfard, Zahra Saaberi, Mostafa Karimi manesh, Seyed Mohammad Hossein Hashemi 외

Research on evaluating and analyzing large language models (LLMs) has been extensive for resource-rich languages such as English, yet their performance in languages such as Persian has received considerably less attentio…

DescriptiveEthicsMultiple-choiceText Generation

PsyEval: A Suite of Mental Health Related Tasks for Evaluating Large Language Models

2023-11-15 · Haoan Jin, Siyuan Chen, Dilawaier Dilixiati, Yewei Jiang 외

Evaluating Large Language Models (LLMs) in the mental health domain poses distinct challenged from other domains, given the subtle and highly subjective nature of symptoms that exhibit significant variability among indiv…

Language ModellingLarge Language ModelModel Optimization