paper-with-me

홈 › Papers

Cross-Model Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Across Three Large Language Models

2026-04-21 · Kihyuk Lee arxiv

This study compared repeated generation consistency of exercise prescription outputs across three large language models (LLMs), specifically GPT-4.1, Claude Sonnet 4.6, and Gemini 2.5 Flash, under temperature=0 conditions. Each model generated prescriptions for six clinical scenarios 20 times, yielding 360 total outputs analyzed across four dimensions: semantic similarity, output reproducibility, FITT classification, and safety expression. Mean semantic similarity was highest for GPT-4.1 (0.955), followed by Gemini 2.5 Flash (0.950) and Claude Sonnet 4.6 (0.903), with significant inter-model differences confirmed (H = 458.41, p < .001). Critically, these scores reflected fundamentally different generative behaviors: GPT-4.1 produced entirely unique outputs (100%) with stable semantic content, while Gemini 2.5 Flash showed pronounced output repetition (27.5% unique outputs), indicating that its high similarity score derived from text duplication rather than consistent reasoning. Identical decoding settings thus yielded fundamentally different consistency profiles, a distinction that single-output evaluations cannot capture. Safety expression reached ceiling levels across all models, confirming its limited utility as a differentiating metric. These results indicate that model selection constitutes a clinical rather than merely technical decision, and that output behavior under repeated generation conditions should be treated as a core criterion for reliable deployment of LLM-based exercise prescription systems.

📄 PDF Abstract BibTeX arXiv:2604.19598

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Similarity

Similar Papers 제목 키워드 기반

Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Using a Large Language Model

2026-04-13 · Kihyuk Lee arxiv

Background: Large language models (LLMs) have been explored as tools for generating personalized exercise prescriptions, yet the consistency of outputs under identical conditions remains insufficiently examined. Objectiv…

Semantic Similarity

Clinician-Directed Large Language Model Software Generation for Therapeutic Interventions in Physical Rehabilitation

2025-11-23 · Edward Kim, Yuri Cho, Jose Eduardo E. Lima, Julie Muccini 외 arxiv

Digital health interventions increasingly deliver home exercise programs via sensor-equipped devices such as smartphones, enabling remote monitoring of adherence and performance. However, current software is usually auth…

Generating Medical Prescriptions with Conditional Transformer

2023-10-30 · Samuel Belkadi, Nicolo Micheletti, Lifeng Han, Warren Del-Pinto 외

Access to real-world medication prescriptions is essential for medical research and healthcare quality improvement. However, access to real medication prescriptions is often limited due to the sensitive nature of the inf…

2kLanguage Modellingnamed-entity-recognitionNamed Entity Recognition+2

Breaking Boundaries: A Chronology with Future Directions of Women in Exercise Physiology Research, Centred on Pregnancy

2024-04-12 · Abbey E. Corson, Meaghan MacDonald, Velislava Tzaneva, Chris M. Edwards 외

Historically, females were excluded from clinical research due to their reproductive roles, hindering medical understanding and healthcare quality. Despite guidelines promoting equal participation, females are underrepre…

Misconceptions

Neuromuscular and Metabolic Responses during Repeated Bouts of Loaded Downhill Walking

2023-09-18 · Emeric Chalchat, Julien Siracusa, Luis Peñailillo, Alexandra Malgoyre 외

The aim of this study was to compare vastus lateralis (VL) and rectus femoris (RF) muscles for their nervous and mechanical adaptations during two bouts of downhill walking (DW) with load carriage performed two weeks apa…