PHEONA: An Evaluation Framework for Large Language Model-based Approaches to Computational Phenotyping
Computational phenotyping is essential for biomedical research but often requires significant time and resources, especially since traditional methods typically involve extensive manual data review. While machine learning and natural language processing advancements have helped, further improvements are needed. Few studies have explored using Large Language Models (LLMs) for these tasks despite known advantages of LLMs for text-based tasks. To facilitate further research in this area, we developed an evaluation framework, Evaluation of PHEnotyping for Observational Health Data (PHEONA), that outlines context-specific considerations. We applied and demonstrated PHEONA on concept classification, a specific task within a broader phenotyping process for Acute Respiratory Failure (ARF) respiratory support therapies. From the sample concepts tested, we achieved high classification accuracy, suggesting the potential for LLM-based methods to improve computational phenotyping processes.
Code (0)
등록된 구현이 없습니다.
Tasks
Computational PhenotypingLanguage ModelingLanguage ModellingLarge Language ModelRespiratory FailureSimilar Papers 제목 키워드 기반
SHREC and PHEONA: Using Large Language Models to Advance Next-Generation Computational Phenotyping
Objective: Computational phenotyping is a central informatics activity with resulting cohorts supporting a wide variety of applications. However, it is time-intensive because of manual data review, limited automation, an…
Computational PhenotypingPrompt EngineeringSpecificitytext-classification+1Multilingual Self-Taught Faithfulness Evaluators
The growing use of large language models (LLMs) has increased the need for automatic evaluation systems, particularly to address the challenge of information hallucination. Although existing faithfulness evaluation appro…
Cross-Lingual TransferMachine TranslationBioACE: An Automated Framework for Biomedical Answer and Citation Evaluations
With the increasing use of large language models (LLMs) for generating answers to biomedical questions, it is crucial to evaluate the quality of the generated answers and the references provided to support the facts in t…
Natural Language InferenceQuestion AnsweringCollabEval: Enhancing LLM-as-a-Judge via Multi-Agent Collaboration
Large Language Models (LLMs) have revolutionized AI-generated content evaluation, with the LLM-as-a-Judge paradigm becoming increasingly popular. However, current single-LLM evaluation approaches face significant challen…
Cross-Lingual LLM-Judge Transfer via Evaluation Decomposition
As large language models are increasingly deployed across diverse real-world applications, extending automated evaluation beyond English has become a critical challenge. Existing evaluation approaches are predominantly E…
Cross-Lingual Transfer