paper-with-me

홈 › Papers

Large Language Models for Automatic Detection of Sensitive Topics

2024-09-02 · Ruoyu Wen, Stephanie Elena Crowe, Kunal Gupta, Xinyue Li, Mark Billinghurst, Simon Hoermann, Dwain Allan, Alaeddin Nassani, Thammathip Piumsomboon

Sensitive information detection is crucial in content moderation to maintain safe online communities. Assisting in this traditionally manual process could relieve human moderators from overwhelming and tedious tasks, allowing them to focus solely on flagged content that may pose potential risks. Rapidly advancing large language models (LLMs) are known for their capability to understand and process natural language and so present a potential solution to support this process. This study explores the capabilities of five LLMs for detecting sensitive messages in the mental well-being domain within two online datasets and assesses their performance in terms of accuracy, precision, recall, F1 scores, and consistency. Our findings indicate that LLMs have the potential to be integrated into the moderation workflow as a convenient and precise detection tool. The best-performing model, GPT-4o, achieved an average accuracy of 99.5\% and an F1-score of 0.99. We discuss the advantages and potential challenges of using LLMs in the moderation workflow and suggest that future research should address the ethical considerations of utilising this technology.

📄 PDF Abstract BibTeX arXiv:2409.00940

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Natural language processing on customer note data

2023-05-03 · Andrew Hilditch, David Webb, Jozef Baca, Tom Armitage 외

Automatic analysis of customer data for businesses is an area that is of interest to companies. Business to business data is studied rarely in academia due to the sensitive nature of such information. Applying natural la…

Keyword ExtractionSentiment Analysis

Hierarchical Latent Semantic Mapping for Automated Topic Generation

2015-11-11 · Guorui Zhou, Guang Chen

Much of information sits in an unprecedented amount of text data. Managing allocation of these large scale text data is an important problem for many areas. Topic modeling performs well in this problem. The traditional g…

Community Detection

Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics

2025-12-18 · Iker García-Ferrero, David Montero, Roman Orus arxiv

We introduce Refusal Steering, an inference-time method to exercise fine-grained control over Large Language Models refusal behaviour on politically sensitive topics without retraining. We replace fragile pattern-based r…

A Survey on Computational Propaganda Detection

2020-07-15 · Giovanni Da San Martino, Stefano Cresci, Alberto Barron-Cedeno, Seunghak Yu 외

Propaganda campaigns aim at influencing people's mindset with the purpose of advancing a specific agenda. They exploit the anonymity of the Internet, the micro-profiling ability of social networks, and the ease of automa…

Propaganda detectionSurvey

Sub-Story Detection in Twitter with Hierarchical Dirichlet Processes

2016-06-11 · Srijith P. K., Hepple Mark, Bontcheva Kalina, Preotiuc-Pietro Daniel

Social media has now become the de facto information source on real world events. The challenge, however, due to the high volume and velocity nature of social media streams, is in how to follow all posts pertaining to a …

Clustering