paper-with-me

Papers

Fusion-Eval: Integrating Assistant Evaluators with LLMs

2023-11-15 · Lei Shu, Nevan Wichers, Liangchen Luo, Yun Zhu, Yinxiao Liu, Jindong Chen, Lei Meng

Evaluating natural language systems poses significant challenges, particularly in the realms of natural language understanding and high-level reasoning. In this paper, we introduce 'Fusion-Eval', an innovative approach that leverages Large Language Models (LLMs) to integrate insights from various assistant evaluators. The LLM is given the example to evaluate along with scores from the assistant evaluators. Each of these evaluators specializes in assessing distinct aspects of responses. Fusion-Eval achieves a 0.962 system-level Kendall-Tau correlation with humans on SummEval and a 0.744 turn-level Spearman correlation on TopicalChat, which is significantly higher than baseline methods. These results highlight Fusion-Eval's significant potential in the realm of natural language system evaluation.

📄 PDF Abstract BibTeX arXiv:2311.09204

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language Understanding

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Collaboration with Conversational AI Assistants for UX Evaluation: Questions and How to Ask them (Voice vs. Text)

2023-03-07 · Emily Kuang, Ehsan Jahangirzadeh Soure, Mingming Fan, Jian Zhao 외

AI is promising in assisting UX evaluators with analyzing usability tests, but its judgments are typically presented as non-interactive visualizations. Evaluators may have questions about test recordings, but have no way…

Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course

2024-07-07 · Cheng-Han Chiang, Wei-Chih Chen, Chun-Yi Kuan, Chienchou Yang 외

Using large language models (LLMs) for automatic evaluation has become an important evaluation method in NLP research. However, it is unclear whether these LLM-based evaluators can be applied in real-world classrooms to …

Instruction FollowingLanguage ModelingLanguage ModellingLarge Language Model

HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

2024-02-24 · Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang 외

Large language models (LLMs) have emerged as a promising alternative to expensive human evaluations. However, the alignment and coverage of LLM-based evaluations are often limited by the scope and potential bias of the e…

Language ModelingLanguage ModellingLarge Language Model

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory

2026-06-02 · Noujoud Nader, Ibrahem Aljabea, Patrick Diehl, Deepti Gupta arxiv

Large language models (LLMs) are increasingly used as self-study assistants in technical disciplines, yet their reliability as mathematical reasoning assistants remains poorly understood. We introduce GTBench, a curricul…

Mathematical Reasoning

Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response

2023-05-24 · Yongkang Liu, Shi Feng, Daling Wang, Yifei Zhang 외

LLMs (large language models) such as ChatGPT have shown remarkable language understanding and generation capabilities. Although reference-free evaluators based on LLMs show better human alignment than traditional referen…

Dialogue Generation