paper-with-me

홈 › Papers

Understanding LLMs in Title-Abstract Screening: From Disagreements to Recommendations

2026-06-16 · Mika Mäntylä, Patricia Matsubara, Katia Romero Felizardo, Miikka Kuutila, Marco Gerosa, Savio de Sousa Sampaio, Tayana Conte, Igor Steinmacher arxiv

Several studies have examined the use of large language models (LLMs) for title-abstract screening in systematic reviews (SRs), reporting mixed accuracy. However, questions of reliability remain largely unaddressed. In this study, we go beyond quantitative LLM-human agreement metrics and qualitatively investigate how and why LLMs fail. We also propose actionable recommendations. We analyzed disagreements between LLMs and researchers across six software engineering SRs and over 1,000 primary study papers. For each SR, papers were screened independently by human experts and LLMs in zero-shot mode, resulting in Kappa values ranging from 0.52 to 0.77. Qualitative analysis suggests that human-LLM disagreement results from recurring, identifiable causes, such as boundary ambiguity in key terms, keyword overemphasization, and incorrect topic inference. Based on these findings, we propose recommendations such as validating semantic understanding before deployment, running multiple LLMs, and focusing validation efforts on borderline cases. Future studies are needed to validate the impact of our recommendations, and community efforts are needed to develop normative guidelines on LLM usage in SRs.

📄 PDF Abstract BibTeX arXiv:2606.17588

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Promise and Challenges of Using LLMs to Accelerate the Screening Process of Systematic Reviews

2024-04-24 · Aleksi Huotala, Miikka Kuutila, Paul Ralph, Mika Mäntylä

Systematic review (SR) is a popular research method in software engineering (SE). However, conducting an SR takes an average of 67 weeks. Thus, automating any step of the SR process could reduce the effort associated wit…

Text Simplification

Leveraging LLMs for Title and Abstract Screening for Systematic Review: A Cost-Effective Dynamic Few-Shot Learning Approach

2025-12-12 · Yun-Chung Liu, Rui Yang, Jonathan Chong Kai Liew, Ziran Yin 외 arxiv

Systematic reviews are a key component of evidence-based medicine, playing a critical role in synthesizing existing research evidence and guiding clinical decisions. However, with the rapid growth of research publication…

Few-Shot Learning

Streamlining Systematic Reviews: A Novel Application of Large Language Models

2024-12-14 · Fouad Trad, Ryan Yammine, Jana Charafeddine, Marlene Chakhtoura 외

Systematic reviews (SRs) are essential for evidence-based guidelines but are often limited by the time-consuming nature of literature screening. We propose and evaluate an in-house system based on Large Language Models (…

ArticlesPrompt EngineeringRAGRetrieval-augmented Generation+1

AISysRev -- LLM-based Tool for Title-abstract Screening

2025-10-08 · Aleksi Huotala, Miikka Kuutila, Olli-Pekka Turtio, Simo Sipilä 외 arxiv

Conducting systematic reviews is laborious. In the screening or study selection phase, the number of papers can be overwhelming. Recent research has demonstrated that large language models (LLMs) can perform title-abstra…

Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs

2026-04-29 · Gilberto Sussumu Hida, Danilo Monteiro Ribeiro, Erika Yahata arxiv

Context: Study screening in systematic literature reviews is costly, inconsistency-prone, and risk-asymmetric, since false negatives can compromise validity. Despite rapid uptake of Large Language Models (LLMs), there is…