paper-with-me

홈 › Papers

Distinct Profiles of Run-to-Run Score Reliability and Expert-Panel Alignment Across Four LLM Evaluators of Simulated Japanese-Language AI-to-AI Counseling

2025-06-28 · Keita Kiuchi, Yoshikazu Fujimoto, Hideyuki Gotō, Tomonori Hosokawa, Makoto Nishimura, Yōsuke Satō, Izumi Sezai, Tomohiro Inoue arxiv

Large language models (LLMs) increasingly evaluate generated dialogue, but repeatable scores do not necessarily align with professional judgment. This observational fixed-benchmark study compared four configured LLM evaluator systems (GPT-5.5, Gemini 3.5 Flash, Claude Opus 4.8, and Fable 5) with aggregated ratings from 15 counseling experts on 18 complete simulated AI-to-AI counseling sessions conducted in Japanese. The sessions represented three counselor conditions across six prespecified client profiles. Each system scored every transcript three times on four motivational interviewing-informed dimensions and overall quality. All four systems assigned higher scores than the expert panel for softening sustain talk and overall quality, although differences varied across systems and constructs. Single-run intraclass correlation coefficients ranged from .33 to .96, showing that high run-to-run reliability did not ensure closer expert-panel alignment. Claude Opus 4.8 had the smallest mean absolute difference, whereas Fable 5 had an intermediate difference. In a secondary benchmark analysis, GPT-4-turbo sessions generated with the Structured Multi-step Dialogue Prompt received higher expert ratings than sessions generated by the same model with a minimal instruction for cultivating change talk, partnership, empathy, and overall quality; the softening sustain talk contrast remained uncertain. The fixed benchmark contained one session per counselor-condition-by-profile cell, so inference concerns these sessions rather than all possible stochastic regenerations. Run-to-run reliability, expert-panel alignment, and condition discrimination are separate properties of automated counseling evaluation. This benchmark supports construct-level assessment of LLM evaluators against professional judgment.

📄 PDF Abstract BibTeX arXiv:2507.02950

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Model as One Rater Among Several: Measuring Political Positions in Data-Sparse Regions with a Language-Model Panel

2026-06-22 · Tarek Gara arxiv

Most tools for measuring political positions, manifesto coding, expert surveys, text-scaling models, were built and validated on Western party systems, and outside that setting they work poorly, and often not at all. Thi…

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

2026-07-20 · Jia-Kai Dong, Yi-Cheng Lin, Hung-yi Lee hf

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching qualit…

Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

2026-04-16 · Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor 외 arxiv

Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators. Here, we evaluate an LLM Jury, composed of three frontier AI…

AI and Open-data Driven Scalable Solar Power Profiling

2026-05-04 · Shiliang Zhang, Sabita Maharjan, Damla Turgut arxiv

Solar photovoltaic (PV) deployment is expanding rapidly, yet detailed, up-to-date information on the spatial distribution and capacity of rooftop PV remains limited. This paper presents an open, scalable framework for de…

The Scientometrics and Reciprocality Underlying Co-Authorship Panels in Google Scholar Profiles

2023-08-14 · Ariel Alexi, Teddy Lazebnik, Ariel Rosenfeld

Online academic profiles are used by scholars to reflect a desired image to their online audience. In Google Scholar, scholars can select a subset of co-authors for presentation in a central location on their profile usi…