paper-with-me

홈 › Papers

CFBench: A Comprehensive Constraints-Following Benchmark for LLMs

2024-08-02 · Chenglin Zhu, Yanjun Shen, Wenjing Luo, Yan Zhang, Hao Liang, Tao Zhang, Fan Yang, MingAn Lin, Yujing Qiao, WeiPeng Chen, Bin Cui, Wentao Zhang, Zenan Zhou

The adeptness of Large Language Models (LLMs) in comprehending and following natural language instructions is critical for their deployment in sophisticated real-world applications. Existing evaluations mainly focus on fragmented constraints or narrow scenarios, but they overlook the comprehensiveness and authenticity of constraints from the user's perspective. To bridge this gap, we propose CFBench, a large-scale Comprehensive Constraints Following Benchmark for LLMs, featuring 1,000 curated samples that cover more than 200 real-life scenarios and over 50 NLP tasks. CFBench meticulously compiles constraints from real-world instructions and constructs an innovative systematic framework for constraint types, which includes 10 primary categories and over 25 subcategories, and ensures each constraint is seamlessly integrated within the instructions. To make certain that the evaluation of LLM outputs aligns with user perceptions, we propose an advanced methodology that integrates multi-dimensional assessment criteria with requirement prioritization, covering various perspectives of constraints, instructions, and requirement fulfillment. Evaluating current leading LLMs on CFBench reveals substantial room for improvement in constraints following, and we further investigate influencing factors and enhancement strategies. The data and code are publicly available at https://github.com/PKU-Baichuan-MLSystemLab/CFBench

📄 PDF Abstract BibTeX arXiv:2408.01122

Code (1)

pku-baichuan-mlsystemlab/cfbench 공식 구현

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

CFBenchmark: Chinese Financial Assistant Benchmark for Large Language Model

2023-11-10 · Yang Lei, Jiangtong Li, Dawei Cheng, Zhijun Ding 외

Large language models (LLMs) have demonstrated great potential in the financial domain. Thus, it becomes important to assess the performance of LLMs in the financial tasks. In this work, we introduce CFBenchmark, to eval…

Language ModelingLanguage ModellingLarge Language Model

CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model

2025-06-16 · Jiangtong Li, Yiyun Zhu, Dawei Cheng, Zhijun Ding 외

Multimodal Large Language Models (MLLMs) have rapidly evolved with the growth of Large Language Models (LLMs) and are now applied in various fields. In finance, the integration of diverse modalities such as text, charts,…

Decision MakingFinancial AnalysisLanguage ModelingLanguage Modelling+2

FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

2023-10-31 · Yuxin Jiang, YuFei Wang, Xingshan Zeng, Wanjun Zhong 외

The ability to follow instructions is crucial for Large Language Models (LLMs) to handle various real-world applications. Existing benchmarks primarily focus on evaluating pure response quality, rather than assessing whe…

Instruction Following

Benchmarking Complex Instruction-Following with Multiple Constraints Composition

2024-07-04 · Bosi Wen, Pei Ke, Xiaotao Gu, Lindong Wu 외

Instruction following is one of the fundamental capabilities of large language models (LLMs). As the ability of LLMs is constantly improving, they have been increasingly applied to deal with complex human instructions in…

BenchmarkingInstruction Following

Benchmarking Large Language Models on Controllable Generation under Diversified Instructions

2024-01-01 · Yihan Chen, Benfeng Xu, Quan Wang, Yi Liu 외

While large language models (LLMs) have exhibited impressive instruction-following capabilities, it is still unclear whether and to what extent they can respond to explicit constraints that might be entailed in various i…

BenchmarkingInstruction FollowingText Generation