paper-with-me

홈 › Papers

Evaluating Large Language Models on a Highly-specialized Topic, Radiation Oncology Physics

2023-04-01 · Jason Holmes, Zhengliang Liu, Lian Zhang, Yuzhen Ding, Terence T. Sio, Lisa A. McGee, Jonathan B. Ashman, Xiang Li, Tianming Liu, Jiajian Shen, Wei Liu

We present the first study to investigate Large Language Models (LLMs) in answering radiation oncology physics questions. Because popular exams like AP Physics, LSAT, and GRE have large test-taker populations and ample test preparation resources in circulation, they may not allow for accurately assessing the true potential of LLMs. This paper proposes evaluating LLMs on a highly-specialized topic, radiation oncology physics, which may be more pertinent to scientific and medical communities in addition to being a valuable benchmark of LLMs. We developed an exam consisting of 100 radiation oncology physics questions based on our expertise at Mayo Clinic. Four LLMs, ChatGPT (GPT-3.5), ChatGPT (GPT-4), Bard (LaMDA), and BLOOMZ, were evaluated against medical physicists and non-experts. ChatGPT (GPT-4) outperformed all other LLMs as well as medical physicists, on average. The performance of ChatGPT (GPT-4) was further improved when prompted to explain first, then answer. ChatGPT (GPT-3.5 and GPT-4) showed a high level of consistency in its answer choices across a number of trials, whether correct or incorrect, a characteristic that was not observed in the human test groups. In evaluating ChatGPTs (GPT-4) deductive reasoning ability using a novel approach (substituting the correct answer with "None of the above choices is the correct answer."), ChatGPT (GPT-4) demonstrated surprising accuracy, suggesting the potential presence of an emergent ability. Finally, although ChatGPT (GPT-4) performed well overall, its intrinsic properties did not allow for further improvement when scoring based on a majority vote across trials. In contrast, a team of medical physicists were able to greatly outperform ChatGPT (GPT-4) using a majority vote. This study suggests a great potential for LLMs to work alongside radiation oncology experts as highly knowledgeable assistants.

📄 PDF Abstract BibTeX arXiv:2304.01938

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음
BLOOMZ BLOOMZ is a Multitask prompted finetuning (MTF) variant of BLOOM.

Similar Papers 제목 키워드 기반

A Comparative Study of Student Perspectives on Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics

2025-12-01 · Suqing Liu, Runlong Ye, Christopher Eaton, Bogdan Simion 외 arxiv

To address the scalability of feedback in computer science while mitigating the privacy and cost limitations of commercial Large Language Models (LLMs), this study evaluates a locally hosted Small Language Model (SLM). W…

PsyEval: A Suite of Mental Health Related Tasks for Evaluating Large Language Models

2023-11-15 · Haoan Jin, Siyuan Chen, Dilawaier Dilixiati, Yewei Jiang 외

Evaluating Large Language Models (LLMs) in the mental health domain poses distinct challenged from other domains, given the subtle and highly subjective nature of symptoms that exhibit significant variability among indiv…

Language ModellingLarge Language ModelModel Optimization

The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead?

2024-10-07 · Alexander S. Choi, Syeda Sabrina Akter, JP Singh, Antonios Anastasopoulos

Large Language Models (LLMs) have shown capabilities close to human performance in various analytical tasks, leading researchers to use them for time and labor-intensive analyses. However, their capability to handle high…

Is Language Modeling Enough? Evaluating Effective Embedding Combinations

2020-05-01 · LREC 2020 5 · Rudolf Schneider, Tom Oberhauser, Paul Grundmann, Felix Alex Gers 외

Universal embeddings, such as BERT or ELMo, are useful for a broad set of natural language processing tasks like text classification or sentiment analysis. Moreover, specialized embeddings also exist for tasks like topic…

Entity DisambiguationGeneral ClassificationLanguage ModelingLanguage Modelling+5

Topic Model or Topic Twaddle? Re-evaluating Semantic Interpretability Measures

2021-06-01 · NAACL 2021 4 · Caitlin Doogan, Wray Buntine

When developing topic models, a critical question that should be asked is: How well will this model work in an applied setting? Because standard performance evaluation of topic interpretability uses automated measures mo…

Topic Models