Mind the data gap(s): Investigating power in speech and language datasets
Algorithmic oppression is an urgent and persistent problem in speech and language technologies. Considering power relations embedded in datasets before compiling or using them to train or test speech and language technologies is essential to designing less harmful, more just technologies. This paper presents a reflective exercise to recognise and challenge gaps and the power relations they reveal in speech and language datasets by applying principles of Data Feminism and Design Justice, and building on work on dataset documentation and sociolinguistics.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors
Reconstructing continuous speech from non-invasive neural recordings is a fundamental problem for probing human auditory perception and building safe, scalable speech brain-computer interfaces. Despite recent progress, i…
EmoTale: An Enacted Speech-emotion Dataset in Danish
While multiple emotional speech corpora exist for commonly spoken languages, there is a lack of functional datasets for smaller (spoken) languages, such as Danish. To our knowledge, Danish Emotional Speech (DES), publish…
Speech Emotion RecognitionThe Interpreter Understands Your Meaning: End-to-end Spoken Language Understanding Aided by Speech Translation
End-to-end spoken language understanding (SLU) remains elusive even with current large pretrained language models on text and speech, especially in multilingual cases. Machine translation has been established as a powerf…
Abstractive Text SummarizationContinual Learningintent-classificationIntent Classification+5Analysis of Emotional Content in Indian Political Speeches
Emotions play an essential role in public speaking. The emotional content of speech has the power to influence minds. As such, we present an analysis of the emotional content of politicians speech in the Indian political…
Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias
Speech Continuation (SC) is the task of generating a coherent extension of a spoken prompt while preserving both semantic context and speaker identity. Because SC is constrained to a single audio stream, it offers a more…