paper-with-me

홈 › Papers

Leveraging LLMs for Semi-Automatic Corpus Filtration in Systematic Literature Reviews

2025-10-13 · Lucas Joos, Daniel A. Keim, Maximilian T. Fischer arxiv

The creation of systematic literature reviews (SLR) is critical for analyzing the landscape of a research field and guiding future research directions. However, retrieving and filtering the literature corpus for an SLR is highly time-consuming and requires extensive manual effort, as keyword-based searches in digital libraries often return numerous irrelevant publications. In this work, we propose a pipeline leveraging multiple large language models (LLMs), classifying papers based on descriptive prompts and deciding jointly using a consensus scheme. The entire process is human-supervised and interactively controlled via our open-source visual analytics web interface, LLMSurver, which enables real-time inspection and modification of model outputs. We evaluate our approach using ground-truth data from a recent SLR comprising over 8,000 candidate papers, benchmarking both open and commercial state-of-the-art LLMs from mid-2024 and fall 2025. Results demonstrate that our pipeline significantly reduces manual effort while achieving lower error rates than single human annotators. Furthermore, modern open-source models prove sufficient for this task, making the method accessible and cost-effective. Overall, our work demonstrates how responsible human-AI collaboration can accelerate and enhance systematic literature reviews within academic workflows.

📄 PDF Abstract BibTeX arXiv:2510.11409

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Creating Domain-Specific Translation Memories for Machine Translation Fine-tuning: The TRENCARD Bilingual Cardiology Corpus

2024-09-04 · Gokhan Dogru

This article investigates how translation memories (TM) can be created by translators or other language professionals in order to compile domain-specific parallel corpora , which can then be used in different scenarios, …

Language ModelingLanguage ModellingLarge Language ModelMachine Translation+1

Cutting Through the Clutter: The Potential of LLMs for Efficient Filtration in Systematic Literature Reviews

2024-07-15 · Lucas Joos, Daniel A. Keim, Maximilian T. Fischer

Systematic literature reviews (SLRs) are essential but labor-intensive due to high publication volumes and inefficient keyword-based filtering. To streamline this process, we evaluate Large Language Models (LLMs) for enh…

Articles

An Extension of the Slovak Broadcast News Corpus based on Semi-Automatic Annotation

2016-05-01 · LREC 2016 5 · Peter Viszlay, J{\'a}n Sta{\v{s}}, Tom{\'a}{\v{s}} Koct{\'u}r, Martin Lojka 외

In this paper, we introduce an extension of our previously released TUKE-BNews-SK corpus based on a semi-automatic annotation scheme. It firstly relies on the automatic transcription of the BN data performed by our Slova…

speech-recognitionSpeech Recognition

A speech corpus for chronic kidney disease

2022-11-03 · Jihyun Mun, Sunhee Kim, Myeong Ju Kim, Jiwon Ryu 외

In this study, we present a speech corpus of patients with chronic kidney disease (CKD) that will be used for research on pathological voice analysis, automatic illness identification, and severity prediction. This paper…

Sentenceseverity prediction

Building a Biomedical Full-Text Part-of-Speech Corpus Semi-Automatically

2022-06-01 · LREC (LAW) 2022 6 · Nicholas Elder, Robert E. Mercer, Sudipta Singha Roy

This paper presents a method for semi-automatically building a corpus of full-text English-language biomedical articles annotated with part-of-speech tags. The outcomes are a semi-automatic procedure to create a large si…

ArticlesTAG