Silent Signals, Loud Impact: LLMs for Word-Sense Disambiguation of Coded Dog Whistles
A dog whistle is a form of coded communication that carries a secondary meaning to specific audiences and is often weaponized for racial and socioeconomic discrimination. Dog whistling historically originated from United States politics, but in recent years has taken root in social media as a means of evading hate speech detection systems and maintaining plausible deniability. In this paper, we present an approach for word-sense disambiguation of dog whistles from standard speech using Large Language Models (LLMs), and leverage this technique to create a dataset of 16,550 high-confidence coded examples of dog whistles used in formal and informal communication. Silent Signals is the largest dataset of disambiguated dog whistle usage, created for applications in hate speech detection, neology, and political science. The dataset can be found at https://huggingface.co/datasets/SALT-NLP/silent_signals.
Code (0)
등록된 구현이 없습니다.
Tasks
Hate Speech DetectionWord Sense DisambiguationSimilar Papers 제목 키워드 기반
Understanding Silent Data Corruption in LLM Training
As the scale of training large language models (LLMs) increases, one emergent failure is silent data corruption (SDC), where hardware produces incorrect computations without explicit failure signals. In this work, we are…
Cloud ComputingDigital Voicing of Silent Speech
In this paper, we consider the task of digitally voicing silent speech, where silently mouthed words are converted to audible speech based on electromyography (EMG) sensor measurements that capture muscle impulses. While…
Electromyography (EMG)Speech SynthesisRead Quietly, Think Aloud: Decoupling Comprehension and Reasoning in LLMs
Large Language Models (LLMs) have demonstrated remarkable proficiency in understanding text and generating high-quality responses. However, a critical distinction from human cognition is their typical lack of a distinct …
From Silent Signals to Natural Language: A Dual-Stage Transformer-LLM Approach
Silent Speech Interfaces (SSIs) have gained attention for their ability to generate intelligible speech from non-acoustic signals. While significant progress has been made in advancing speech generation pipelines, limite…
Speech RecognitionLoud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is routine. Two questions then matter beyond t…