paper-with-me

Papers

Facilitating Holistic Evaluations with LLMs: Insights from Scenario-Based Experiments

2024-05-28 · Toru Ishida, Tongxi Liu, Hailong Wang, William K. Cheunga

Workshop courses designed to foster creativity are gaining popularity. However, even experienced faculty teams find it challenging to realize a holistic evaluation that accommodates diverse perspectives. Adequate deliberation is essential to integrate varied assessments, but faculty often lack the time for such exchanges. Deriving an average score without discussion undermines the purpose of a holistic evaluation. Therefore, this paper explores the use of a Large Language Model (LLM) as a facilitator to integrate diverse faculty assessments. Scenario-based experiments were conducted to determine if the LLM could integrate diverse evaluations and explain the underlying pedagogical theories to faculty. The results were noteworthy, showing that the LLM can effectively facilitate faculty discussions. Additionally, the LLM demonstrated the capability to create evaluation criteria by generalizing a single scenario-based experiment, leveraging its already acquired pedagogical domain knowledge.

📄 PDF Abstract BibTeX arXiv:2405.17728

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

$\texttt{SAGE}$: A Generic Framework for LLM Safety Evaluation

2025-04-28 · Madhur Jindal, Hari Shrawgi, Parag Agrawal, Sandipan Dandapat

Safety evaluation of Large Language Models (LLMs) has made progress and attracted academic interest, but it remains challenging to keep pace with the rapid integration of LLMs across diverse applications. Different appli…

Red TeamingSafety Alignment

Re3: A Holistic Framework and Dataset for Modeling Collaborative Document Revision

2024-05-31 · Qian Ruan, Ilia Kuznetsov, Iryna Gurevych

Collaborative review and revision of textual documents is the core of knowledge work and a promising target for empirical analysis and NLP assistance. Yet, a holistic framework that would allow modeling complex relations…

ChEF: A Comprehensive Evaluation Framework for Standardized Assessment of Multimodal Large Language Models

2023-11-05 · Zhelun Shi, Zhipin Wang, Hongxing Fan, Zhenfei Yin 외

Multimodal Large Language Models (MLLMs) have shown impressive abilities in interacting with visual content with myriad potential downstream tasks. However, even though a list of benchmarks has been proposed, the capabil…

HallucinationIn-Context LearningInstruction FollowingQuestion Answering

Enterprise Large Language Model Evaluation Benchmark

2025-06-25 · Liya Wang, David Yi, Damien Jose, John Passarelli 외

Large Language Models (LLMs) ) have demonstrated promise in boosting productivity across AI-powered tools, yet existing benchmarks like Massive Multitask Language Understanding (MMLU) inadequately assess enterprise-speci…

Language Model EvaluationLanguage ModelingLanguage ModellingLarge Language Model+4

An Automated Explainable Educational Assessment System Built on LLMs

2024-12-17 · Jiazheng Li, Artem Bobrov, David West, Cesare Aloisi 외

In this demo, we present AERA Chat, an automated and explainable educational assessment system designed for interactive and visual evaluations of student responses. This system leverages large language models (LLMs) to g…