paper-with-me

홈 › Papers

Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts

2026-04-24 · Zhyar Rzgar K. Rostam, Márta Péntek, János Tibor Czere, Zsombor Zrubka, László Gulácsi, Gábor Kertész arxiv

The rapid increase in scientific publications leads to the fact that manual study screening in systematic literature reviews (SLRs) is increasingly resource consuming, inefficient, and inconsistent. Classifying studies that clearly report health-related quality-of-life results, such as EQ-5D data, requires a high level of clinical interpretation and poses challenges for human reviewers. This study investigates the use of Google's Gemini and Gemma large language models (LLMs) in automating EQ-5D detection in the PubMed biomedical database based only on published abstracts. A multi-phase framework is proposed that integrates few-shot prompting, weight ensembling aggregation, and a soft stacking meta-classifier. Nine LLMs are evaluated on a dataset of PubMed studies manually labeled by two experts regarding EQ-5D reporting. The weighted ensemble of gemini-2.5-pro, gemma-3-12b, and gemma-3-27b obtained a 0.74 weighted F1-score and 0.74 accuracy, exceeding individually attained results. The ensembling of top-performing models improved the balance between precision and recall compared to individual models, while the soft stacking approach provided greater reliability and interpretability. Feature analysis shows that the probability results from the models are important in guiding the final predictions. The findings suggest that an ensemble-based LLM setup is a reliable and scalable approach for automating screening in biomedical research.

📄 PDF Abstract BibTeX arXiv:2606.19345

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CU-UD: text-mining drug and chemical-protein interactions with ensembles of BERT-based models

2021-11-11 · Mehmet Efruz Karabulut, K. Vijay-Shanker, Yifan Peng

Identifying the relations between chemicals and proteins is an important text mining task. BioCreative VII track 1 DrugProt task aims to promote the development and evaluation of systems that can automatically detect rel…

DrugProt

Detecting Causal Language Use in Science Findings

2019-11-01 · IJCNLP 2019 11 · Bei Yu, Yingya Li, Jun Wang

Causal interpretation of correlational findings from observational studies has been a major type of misinformation in science communication. Prior studies on identifying inappropriate use of causal language relied on man…

MisinformationPrediction

ProtoBERT-LoRA: Parameter-Efficient Prototypical Finetuning for Immunotherapy Study Identification

2025-03-26 · Shijia Zhang, Xiyu Ding, Kai Ding, Jacob Zhang 외

Identifying immune checkpoint inhibitor (ICI) studies in genomic repositories like Gene Expression Omnibus (GEO) is vital for cancer research yet remains challenging due to semantic ambiguity, extreme class imbalance, an…

BRAINCELL-AID: An Agentic AI Created Brain Cell Type Resource for Community Annotation

2025-10-20 · Rongbin Li, Wenbo Chen, Zhao Li, Rodrigo Munoz-Castaneda 외 arxiv

Single-cell RNA sequencing has transformed our ability to identify diverse cell types and their transcriptomic signatures. However, annotating these signatures-especially those involving poorly characterized genes-remain…

Assigning function to protein-protein interactions: a weakly supervised BioBERT based approach using PubMed abstracts

2020-08-20 · Aparna Elangovan, Melissa Davis, Karin Verspoor

Motivation: Protein-protein interactions (PPI) are critical to the function of proteins in both normal and diseased cells, and many critical protein functions are mediated by interactions.Knowledge of the nature of these…