paper-with-me

홈 › Papers

DEVAL: A Framework for Evaluating and Improving the Derivation Capability of Large Language Models

2025-11-18 · Yifan Li, Qin Li, Min Zhang, Min Zhang arxiv

Assessing the reasoning ability of Large Language Models (LLMs) over data remains an open and pressing research question. Compared with LLMs, human reasoning can derive corresponding modifications to the output based on certain kinds of changes to the input. This reasoning pattern, which relies on abstract rules that govern relationships between changes of data, has not been comprehensively described or evaluated in LLMs. In this paper, we formally define this reasoning pattern as the Derivation Relation (DR) and introduce the concept of Derivation Capability (DC), i.e. applying DR by making the corresponding modification to the output whenever the input takes certain changes. To assess DC, a systematically constructed evaluation framework named DEVAL is proposed and used to evaluate five popular LLMs and one Large Reasoning Model in seven mainstream tasks. The evaluation results show that mainstream LLMs, such as GPT-4o and Claude3.5, exhibit moderate DR recognition capabilities but reveal significant drop-offs on applying DR effectively in problem-solving scenarios. To improve this, we propose a novel prompt engineering approach called Derivation Prompting (DP). It achieves an average improvement of 15.2% in DC for all tested LLMs, outperforming commonly used prompt engineering techniques.

📄 PDF Abstract BibTeX arXiv:2511.14813

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Similar Papers 제목 키워드 기반

PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation

2025-10-01 · Yujia Xiao, Liumeng Xue, Lei He, Xinyi Chen 외 arxiv

Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into assessing generative capabilities remain…

Audio Generation

SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models

2025-08-08 · Hanqing Wang, Yuan Tian, Mingyu Liu, Zhenhao Zhang 외 arxiv

In the rapidly evolving landscape of Multimodal Large Language Models (MLLMs), the safety concerns of their outputs have earned significant attention. Although numerous datasets have been proposed, they may become outdat…

FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom

2024-04-18 · Yuanqin He, Yan Kang, Lixin Fan, Qiang Yang

Federated Learning (FL) has emerged as a promising solution for collaborative training of large language models (LLMs). However, the integration of LLMs into FL introduces new challenges, particularly concerning the eval…

Federated LearningPrivacy Preserving

RadEval: A framework for radiology text evaluation

2025-09-22 · Justin Xu, Xi Zhang, Javid Abderezaei, Julie Bauml 외 arxiv

We introduce RadEval, a unified, open-source framework for evaluating radiology texts. RadEval consolidates a diverse range of metrics, from classic n-gram overlap (BLEU, ROUGE) and contextual measures (BERTScore) to cli…

MDEval: Evaluating and Enhancing Markdown Awareness in Large Language Models

2025-01-25 · Zhongpu Chen, Yinfeng Liu, Long Shi, Zhi-Jie Wang 외

Large language models (LLMs) are expected to offer structured Markdown responses for the sake of readability in web chatbots (e.g., ChatGPT). Although there are a myriad of metrics to evaluate LLMs, they fail to evaluate…