Distinct Profiles of Run-to-Run Score Reliability and Expert-Panel Alignment Across Four LLM Evaluators of Simulated Japanese-Language AI-to-AI Counseling
Large language models (LLMs) increasingly evaluate generated dialogue, but repeatable scores do not necessarily align with professional judgment. This observational fixed-benchmark study compared four configured LLM evaluator systems (GPT-5.5, Gemini 3.5 Flash, Claude Opus 4.8, and Fable 5) with aggregated ratings from 15 counseling experts on 18 complete simulated AI-to-AI counseling sessions conducted in Japanese. The sessions represented three counselor conditions across six prespecified client profiles. Each system scored every transcript three times on four motivational interviewing-informed dimensions and overall quality. All four systems assigned higher scores than the expert panel for softening sustain talk and overall quality, although differences varied across systems and constructs. Single-run intraclass correlation coefficients ranged from .33 to .96, showing that high run-to-run reliability did not ensure closer expert-panel alignment. Claude Opus 4.8 had the smallest mean absolute difference, whereas Fable 5 had an intermediate difference. In a secondary benchmark analysis, GPT-4-turbo sessions generated with the Structured Multi-step Dialogue Prompt received higher expert ratings than sessions generated by the same model with a minimal instruction for cultivating change talk, partnership, empathy, and overall quality; the softening sustain talk contrast remained uncertain. The fixed benchmark contained one session per counselor-condition-by-profile cell, so inference concerns these sessions rather than all possible stochastic regenerations. Run-to-run reliability, expert-panel alignment, and condition discrimination are separate properties of automated counseling evaluation. This benchmark supports construct-level assessment of LLM evaluators against professional judgment.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
The Model as One Rater Among Several: Measuring Political Positions in Data-Sparse Regions with a Language-Model Panel
Most tools for measuring political positions, manifesto coding, expert surveys, text-scaling models, were built and validated on Western party systems, and outside that setting they work poorly, and often not at all. Thi…
EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching qualit…
Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators. Here, we evaluate an LLM Jury, composed of three frontier AI…
AI and Open-data Driven Scalable Solar Power Profiling
Solar photovoltaic (PV) deployment is expanding rapidly, yet detailed, up-to-date information on the spatial distribution and capacity of rooftop PV remains limited. This paper presents an open, scalable framework for de…
The Scientometrics and Reciprocality Underlying Co-Authorship Panels in Google Scholar Profiles
Online academic profiles are used by scholars to reflect a desired image to their online audience. In Google Scholar, scholars can select a subset of co-authors for presentation in a central location on their profile usi…