paper-with-me

Papers

Objective Metrics for Evaluating Large Language Models Using External Data Sources

2025-08-01 · Haoze Du, Richard Li, Edward Gehringer arxiv

Evaluating the performance of Large Language Models (LLMs) is a critical yet challenging task, particularly when aiming to avoid subjective assessments. This paper proposes a framework for leveraging subjective metrics derived from the class textual materials across different semesters to assess LLM outputs across various tasks. By utilizing well-defined benchmarks, factual datasets, and structured evaluation pipelines, the approach ensures consistent, reproducible, and bias-minimized measurements. The framework emphasizes automation and transparency in scoring, reducing reliance on human interpretation while ensuring alignment with real-world applications. This method addresses the limitations of subjective evaluation methods, providing a scalable solution for performance assessment in educational, scientific, and other high-stakes domains.

📄 PDF Abstract BibTeX arXiv:2508.08277

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DrawingBench: Evaluating Spatial Reasoning and UI Interaction Capabilities of Large Language Models through Mouse-Based Drawing Tasks

2025-12-01 · Hyunjun Kim, Sooyoung Ryu arxiv

As agentic AI systems increasingly operate autonomously, establishing trust through verifiable evaluation becomes critical. Yet existing benchmarks lack the transparency and auditability needed to assess whether agents b…

Spatial Reasoning

EvalCrafter: Benchmarking and Evaluating Large Video Generation Models

2023-10-17 · CVPR 2024 1 · Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang 외

The vision and language generative models have been overgrown in recent years. For video generation, various open-sourced models and public-available services have been developed to generate high-quality videos. However,…

BenchmarkingLanguage ModellingLarge Language ModelText-to-Video Generation+2

Natural Language Generation Using Reinforcement Learning with External Rewards

2019-11-26 · Vidhushini Srinivasan, Sashank Santhanam, Samira Shaikh

We propose an approach towards natural language generation using a bidirectional encoder-decoder which incorporates external rewards through reinforcement learning (RL). We use attention mechanism and maximum mutual info…

Decoderreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

Self-Refinement of Language Models from External Proxy Metrics Feedback

2024-02-27 · Keshav Ramji, Young-suk Lee, Ramón Fernandez Astudillo, Md Arafat Sultan 외

It is often desirable for Large Language Models (LLMs) to capture multiple objectives when providing a response. In document-grounded response generation, for example, agent responses are expected to be relevant to a use…

Question AnsweringResponse Generation

AI Psychometrics: Evaluating the Psychological Reasoning of Large Language Models with Psychometric Validities

2026-03-11 · Yibai Li, Xiaolin Lin, Zhenghui Sha, Zhiye Jin 외 arxiv

The immense number of parameters and deep neural networks make large language models (LLMs) rival the complexity of human brains, which also makes them opaque ``black box'' systems that are challenging to evaluate and in…