paper-with-me

Papers

C$^2$LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation

2024-12-06 · Yanyang Li, Tin Long Wong, Cheung To Hung, Jianqiao Zhao, Duo Zheng, Ka Wai Liu, Michael R. Lyu, LiWei Wang

Recent advances in large language models (LLMs) have shown significant promise, yet their evaluation raises concerns, particularly regarding data contamination due to the lack of access to proprietary training data. To address this issue, we present C$^2$LEVA, a comprehensive bilingual benchmark featuring systematic contamination prevention. C$^2$LEVA firstly offers a holistic evaluation encompassing 22 tasks, each targeting a specific application or ability of LLMs, and secondly a trustworthy assessment due to our contamination-free tasks, ensured by a systematic contamination prevention strategy that fully automates test data renewal and enforces data protection during benchmark data release. Our large-scale evaluation of 15 open-source and proprietary models demonstrates the effectiveness of C$^2$LEVA.

📄 PDF Abstract BibTeX arXiv:2412.04947

Code (1)

lavi-lab/cleva

Tasks

Language Model EvaluationLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

CLEVA: Chinese Language Models EVAluation Platform

2023-08-09 · Yanyang Li, Jianqiao Zhao, Duo Zheng, Zi-Yuan Hu 외

With the continuous emergence of Chinese Large Language Models (LLMs), how to evaluate a model's capabilities has become an increasingly significant issue. The absence of a comprehensive Chinese benchmark that thoroughly…

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

2024-03-12 · Naman jain, King Han, Alex Gu, Wen-Ding Li 외

Large Language Models (LLMs) applied to code-related applications have emerged as a prominent field, attracting significant interest from both academia and industry. However, as new and improved LLMs are developed, exist…

Code GenerationHumanEvalmbpp

Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling

2025-09-30 · Seiji Maekawa, Jackson Hassell, Pouya Pezeshkpour, Tom Mitchell 외 arxiv

Existing benchmarks for tool-augmented language models (TaLMs) lack fine-grained control over task difficulty and remain vulnerable to data contamination. We present FuncBenchGen, a unified, contamination-free framework …

CODECLEANER: Elevating Standards with A Robust Data Contamination Mitigation Toolkit

2024-11-16 · Jialun Cao, Songqiang Chen, Wuqi Zhang, Hau Ching Lo 외

Data contamination presents a critical barrier preventing widespread industrial adoption of advanced software engineering techniques that leverage code language models (CLMs). This phenomenon occurs when evaluation data …

Investigating Data Contamination for Pre-training Language Models

2024-01-11 · Minhao Jiang, Ken Ziyu Liu, Ming Zhong, Rylan Schaeffer 외

Language models pre-trained on web-scale corpora demonstrate impressive capabilities on diverse downstream tasks. However, there is increasing concern whether such capabilities might arise from evaluation datasets being …

Language ModelingLanguage Modelling