Semi-Supervised Language Models for Identification of Personal Health Experiential from Twitter Data: A Case for Medication Effects
First-hand experience related to any changes of one’s health condition and understanding such experience can play an important role in advancing medical science and healthcare. Monitoring the safe use of medication drugs is an important task of pharmacovigilance, and first-hand experience of effects about consumers’ medication intake can be valuable to gain insight into how our human body reacts to medications. Social media have been considered as a possible alternative data source for gathering personal experience with medications posted by users. Identifying personal experience tweets is a challenging classification task, and efforts have made to tackle the challenges using supervised approaches requiring annotated data. There exists abundance of unlabeled Twitter data, and being able to use such data for training without suffering in classification performance is of great value, which can reduce the cost of laborious annotation process. We investigated two semi-supervised learning methods, with different mixes of labeled and unlabeled data in the training set, to understand the impact on classification performance. Our results from both pseudo-label and consistency regularization methods show that both methods generated a noticeable improvement in F1 score when the labeled set was small, and consistency regularization could still provide a small gain even a larger labeled set was used.
Code (0)
등록된 구현이 없습니다.
Tasks
PharmacovigilancePseudo LabelSimilar Papers 제목 키워드 기반
A Semi-supervised Approach for De-identification of Swedish Clinical Text
An abundance of electronic health records (EHR) is produced every day within healthcare. The records possess valuable information for research and future improvement of healthcare. Multiple efforts have been done to prot…
De-identificationAspirinSum: an Aspect-based utility-preserved de-identification Summarization framework
Due to the rapid advancement of Large Language Model (LLM), the whole community eagerly consumes any available text data in order to train the LLM. Currently, large portion of the available text data are collected from i…
De-identificationLanguage ModellingLarge Language ModelSentencePersonalized Anomaly Detection in PPG Data using Representation Learning and Biometric Identification
Photoplethysmography (PPG) signals, typically acquired from wearable devices, hold significant potential for continuous fitness-health monitoring. In particular, heart conditions that manifest in rare and subtle deviatin…
Anomaly DetectionPhotoplethysmography (PPG)Representation LearningUnsupervised Anomaly DetectionA Semisupervised Approach for Language Identification based on Ladder Networks
In this study we address the problem of training a neuralnetwork for language identification using both labeled and unlabeled speech samples in the form of i-vectors. We propose a neural network architecture that can als…
DenoisingLanguage IdentificationSemi-Supervised Graph Representation Learning with Human-centric Explanation for Predicting Fatty Liver Disease
Addressing the challenge of limited labeled data in clinical settings, particularly in the prediction of fatty liver disease, this study explores the potential of graph representation learning within a semi-supervised le…
Feature ImportanceGraph Representation LearningRepresentation Learning