Utilizing Large Language Models to Identify Reddit Users Considering Vaping Cessation for Digital Interventions
The widespread adoption of social media platforms globally not only enhances users' connectivity and communication but also emerges as a vital channel for the dissemination of health-related information, thereby establishing social media data as an invaluable organic data resource for public health research. The surge in popularity of vaping or e-cigarette use in the United States and other countries has caused an outbreak of e-cigarette and vaping use-associated lung injury (EVALI), leading to hospitalizations and fatalities in 2019, highlighting the urgency to comprehend vaping behaviors and develop effective strategies for cession. In this study, we extracted a sample dataset from one vaping sub-community on Reddit to analyze users' quit vaping intentions. Leveraging large language models including both the latest GPT-4 and traditional BERT-based language models for sentence-level quit-vaping intention prediction tasks, this study compares the outcomes of these models against human annotations. Notably, when compared to human evaluators, GPT-4 model demonstrates superior consistency in adhering to annotation guidelines and processes, showcasing advanced capabilities to detect nuanced user quit-vaping intentions that human evaluators might overlook. These preliminary findings emphasize the potential of GPT-4 in enhancing the accuracy and reliability of social media data analysis, especially in identifying subtle users' intentions that may elude human detection.
Code (0)
등록된 구현이 없습니다.
Tasks
Human DetectionSentenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Detecting Reddit Users with Depression Using a Hybrid Neural Network SBERT-CNN
Depression is a widespread mental health issue, affecting an estimated 3.8% of the global population. It is also one of the main contributors to disability worldwide. Recently it is becoming popular for individuals to us…
Sentencetext-classificationText ClassificationRedDust: a Large Reusable Dataset of Reddit User Traits
Social media is a rich source of assertions about personal traits, such as {``}I am a doctor{''} or {``}my hobby is playing tennis{''}. Precisely identifying explicit assertions is difficult, though, because of the users…
AttributeLinguistic Analysis of Schizophrenia in Reddit Posts
We explore linguistic indicators of schizophrenia in Reddit discussion forums. Schizophrenia (SZ) is a chronic mental disorder that affects a person{'}s thoughts and behaviors. Identifying and detecting signs of SZ is di…
BIG-bench Machine LearningCLPsych 2019 Shared Task: Predicting the Degree of Suicide Risk in Reddit Posts
The shared task for the 2019 Workshop on Computational Linguistics and Clinical Psychology (CLPsych{'}19) introduced an assessment of suicide risk based on social media postings, using data from Reddit to identify users …
Analyzing User Perceptions of Large Language Models (LLMs) on Reddit: Sentiment and Topic Modeling of ChatGPT and DeepSeek Discussions
While there is an increased discourse on large language models (LLMs) like ChatGPT and DeepSeek, there is no comprehensive understanding of how users of online platforms, like Reddit, perceive these models. This is an im…