A Federated Learning Approach to Privacy Preserving Offensive Language Identification
The spread of various forms of offensive speech online is an important concern in social media. While platforms have been investing heavily in ways of coping with this problem, the question of privacy remains largely unaddressed. Models trained to detect offensive language on social media are trained and/or fine-tuned using large amounts of data often stored in centralized servers. Since most social media data originates from end users, we propose a privacy preserving decentralized architecture for identifying offensive language online by introducing Federated Learning (FL) in the context of offensive language identification. FL is a decentralized architecture that allows multiple models to be trained locally without the need for data sharing hence preserving users' privacy. We propose a model fusion approach to perform FL. We trained multiple deep learning models on four publicly available English benchmark datasets (AHSD, HASOC, HateXplain, OLID) and evaluated their performance in detail. We also present initial cross-lingual experiments in English and Spanish. We show that the proposed model fusion approach outperforms baselines in all the datasets while preserving privacy.
Code (0)
등록된 구현이 없습니다.
Tasks
Federated LearningLanguage IdentificationPrivacy PreservingSimilar Papers 제목 키워드 기반
A Federated Approach to Few-Shot Hate Speech Detection for Marginalized Communities
Hate speech online remains an understudied issue for marginalized communities, and has seen rising relevance, especially in the Global South, which includes developing societies with increasing internet penetration. In t…
Federated LearningFew-Shot LearningHate Speech DetectionPrivacy PreservingNLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey
Natural Language Processing (NLP) is integral to social media analytics but often processes content containing Personally Identifiable Information (PII), behavioral cues, and metadata raising privacy risks such as survei…
Native Language IdentificationSentiment AnalysisPrivacy Preserving Machine Learning Workflow: from Anonymization to Personalized Differential Privacy Budgets in Federated Learning
The growing development of artificial intelligence based solutions, together with privacy legislation, has driven the rise of the so-called privacy preserving machine learning architectures, such as federated learning. W…
Federated LearningDPCOVID: Privacy-Preserving Federated Covid-19 Detection
Coronavirus (COVID-19) has shown an unprecedented global crisis by the detrimental effect on the global economy and health. The number of COVID-19 cases has been rapidly increasing, and there is no sign of stopping. It l…
Federated LearningPrivacy PreservingFedVLN: Privacy-preserving Federated Vision-and-Language Navigation
Data privacy is a central problem for embodied agents that can perceive the environment, communicate with humans, and act in the real world. While helping humans complete tasks, the agent may observe and process sensitiv…
Privacy PreservingVision and Language Navigation