paper-with-me

홈 › Papers

Ensemble of Large Language Models for Curated Labeling and Rating of Free-text Data

2025-01-14 · Jiaxing Qiu, Dongliang Guo, Papini Natalie, Peace Noelle, Levinson Cheri, Teague R. Henry

Free-text responses are commonly collected in psychological studies, providing rich qualitative insights that quantitative measures may not capture. Labeling curated topics of research interest in free-text data by multiple trained human coders is typically labor-intensive and time-consuming. Though large language models (LLMs) excel in language processing, LLM-assisted labeling techniques relying on closed-source LLMs cannot be directly applied to free-text data, without explicit consent for external use. In this study, we propose a framework of assembling locally-deployable LLMs to enhance the labeling of predetermined topics in free-text data under privacy constraints. Analogous to annotation by multiple human raters, this framework leverages the heterogeneity of diverse open-source LLMs. The ensemble approach seeks a balance between the agreement and disagreement across LLMs, guided by a relevancy scoring methodology that utilizes embedding distances between topic descriptions and LLMs' reasoning. We evaluated the ensemble approach using both publicly accessible Reddit data from eating disorder related forums, and free-text responses from eating disorder patients, both complemented by human annotations. We found that: (1) there is heterogeneity in the performance of labeling among same-sized LLMs, with some showing low sensitivity but high precision, while others exhibit high sensitivity but low precision. (2) Compared to individual LLMs, the ensemble of LLMs achieved the highest accuracy and optimal precision-sensitivity trade-off in predicting human annotations. (3) The relevancy scores across LLMs showed greater agreement than dichotomous labels, indicating that the relevancy scoring method effectively mitigates the heterogeneity in LLMs' labeling.

📄 PDF Abstract BibTeX arXiv:2501.08413

Code (1)

JiaxingQiu/llm_topic_extraction 공식 구현

Tasks

Sensitivity

Similar Papers 제목 키워드 기반

Large language models enabled multiagent ensemble method for efficient EHR data labeling

2024-10-21 · Jingwei Huang, Kuroush Nezafati, Ismael Villanueva-Miranda, Zifan Gu 외

This study introduces a novel multiagent ensemble method powered by LLMs to address a key challenge in ML - data labeling, particularly in large-scale EHR datasets. Manual labeling of such datasets requires domain expert…

Hallucination

Effective Proxy for Human Labeling: Ensemble Disagreement Scores in Large Language Models for Industrial NLP

2023-09-11 · Wei Du, Laksh Advani, Yashmeet Gambhir, Daniel J Perry 외

Large language models (LLMs) have demonstrated significant capability to generalize across a large number of NLP tasks. For industry applications, it is imperative to assess the performance of the LLM on unlabeled produc…

Keyphrase Extraction

Alignment based Sequence Ensemble with Multiple Results from a Single Neural Model Architecture

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Sequence labeling is a fundamental framework that provides the elemental structure and content information for additional natural language processing. However, existing proposed ensemble approaches do not focus on sequen…

POS

Harnessing Multiple Large Language Models: A Survey on LLM Ensemble

2025-02-25 · Zhijun Chen, Jingzheng Li, Pengpeng Chen, Zhuoran Li 외

LLM Ensemble -- which involves the comprehensive use of multiple large language models (LLMs), each aimed at handling user queries during downstream inference, to benefit from their individual strengths -- has gained sub…

Survey

Sequence Alignment Ensemble with a Single Neural Network for Sequence Labeling

2022-07-07 · IEEE Access 2022 7 · Jeesu Jung, SangKeun Jung, Hyein Seo, Hyuk Namgung 외

Sequence labeling, in which a class or label is assigned to each token in a given input order, is a fundamental task in natural language processing. Many advanced neural network architectures have recently been proposed …

Part-Of-Speech TaggingPOSPOS Tagging