paper-with-me

Papers

AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models

2025-01-23 · Yang Fan

As Large Language Models (LLMs) are pretrained on massive-scale corpora, the issue of data contamination has become increasingly severe, leading to potential overestimation of model performance during evaluation. To address this, we propose AdEval (Alignment-based Dynamic Evaluation), a dynamic data evaluation method aimed at mitigating the impact of data contamination on evaluation reliability. AdEval extracts key knowledge points and main ideas to align dynamically generated questions with static data's core concepts. It also leverages online search to provide detailed explanations of related knowledge points, thereby creating high-quality evaluation samples with robust knowledge support. Furthermore, AdEval incorporates mechanisms to control the number and complexity of questions, enabling dynamic alignment and flexible adjustment. This ensures that the generated questions align with the complexity of static data while supporting varied complexity levels. Based on Bloom's taxonomy, AdEval conducts a multi-dimensional evaluation of LLMs across six cognitive levels: remembering, understanding, applying, analyzing, evaluating, and creating. Experimental results on multiple datasets demonstrate that AdEval effectively reduces the impact of data contamination on evaluation outcomes, enhancing both the fairness and reliability of the evaluation process.

📄 PDF Abstract BibTeX arXiv:2501.13983

Code (0)

등록된 구현이 없습니다.

Tasks

Fairness

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

RadEval: A framework for radiology text evaluation

2025-09-22 · Justin Xu, Xi Zhang, Javid Abderezaei, Julie Bauml 외 arxiv

We introduce RadEval, a unified, open-source framework for evaluating radiology texts. RadEval consolidates a diverse range of metrics, from classic n-gram overlap (BLEU, ROUGE) and contextual measures (BERTScore) to cli…

RadOT-Eval: Auditable Structured-Evidence Transport for Radiology Report Evaluation

2026-06-07 · Weixin Liu, Juming Xiong, Yang Li, Qingyuan Song 외 arxiv

Automatic evaluation is critical for high-stakes text generation, where errors often involve omitted findings, hallucinated content, polarity reversals, location changes, uncertainty mismatches, and temporal-comparison e…

Text Generation

VERT: Reliable LLM Judges for Radiology Report Evaluation

2026-04-03 · Federica Bologna, Jean-Philippe Corbeil, Matthew Wilkens, Asma Ben Abacha arxiv

Current literature on radiology report evaluation has focused primarily on designing LLM-based metrics and fine-tuning small models for chest X-rays. However, it remains unclear whether these approaches are robust when a…

parameter-efficient fine-tuning

GEMA-Score: Granular Explainable Multi-Agent Score for Radiology Report Evaluation

2025-03-07 · Zhenxuan Zhang, Kinhei Lee, Weihang Deng, Huichi Zhou 외

Automatic medical report generation supports clinical diagnosis, reduces the workload of radiologists, and holds the promise of improving diagnosis consistency. However, existing evaluation metrics primarily assess the a…

Large Language ModelMedical Report GenerationNER

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

2025-11-16 · Zhanheng Nie, Chenghan Fu, Daoze Zhang, Junxian Wu 외 arxiv

Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challenges: (i) the modality imbalance induced by modality mixed training; (ii)…

Representation Learning