paper-with-me

홈 › Papers

Beyond Benchmarking: A New Paradigm for Evaluation and Assessment of Large Language Models

2024-07-10 · Jin Liu, Qingquan Li, Wenlong Du

In current benchmarks for evaluating large language models (LLMs), there are issues such as evaluation content restriction, untimely updates, and lack of optimization guidance. In this paper, we propose a new paradigm for the measurement of LLMs: Benchmarking-Evaluation-Assessment. Our paradigm shifts the "location" of LLM evaluation from the "examination room" to the "hospital". Through conducting a "physical examination" on LLMs, it utilizes specific task-solving as the evaluation content, performs deep attribution of existing problems within LLMs, and provides recommendation for optimization.

📄 PDF Abstract BibTeX arXiv:2407.07531

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark

2026-06-09 · Hongyu Jin, Siyi Wang, Yang Xiao, Jiaheng Dong 외 arxiv

Humans process rich auditory environments through tightly integrated cognitive capabilities such as audio perception, audio reasoning, and memory. Despite recent progress in large audio-language models (LALMs) across spe…

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform

2026-05-18 · Yu Shang, Yinzhou Tang, Yiding Ma, Zhuohang Li 외 arxiv

World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, existing embodied world model benchmarks a…

BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding

2026-08-04 · Yangxuan Zhou, Yuning Chen, Chen Wu, Jiquan Wang 외 arxiv

Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language instructions, signal processing, quantitative evidence, and scientific int…

Beyond single-channel agentic benchmarking

2026-02-05 · Nelu D. Radpour arxiv

Contemporary benchmarks for agentic artificial intelligence (AI) frequently evaluate safety through isolated task-level accuracy thresholds, implicitly treating autonomous systems as single points of failure. This single…

A Progressive Visual Analytics Tool for Incremental Experimental Evaluation

2019-04-18 · Fabio Giachelle, Gianmaria Silvello

This paper presents a visual tool, AVIATOR, that integrates the progressive visual analytics paradigm in the IR evaluation process. This tool serves to speed-up and facilitate the performance assessment of retrieval mode…

Retrieval