Just Because We Camp, Doesn't Mean We Should: The Ethics of Modelling Queer Voices
Modern voice cloning models claim to be able to capture a diverse range of voices. We test the ability of a typical pipeline to capture the style known colloquially as "gay voice" and notice a homogenisation effect: synthesised speech is rated as sounding significantly "less gay" (by LGBTQ+ participants) than its corresponding ground-truth for speakers with "gay voice", but ratings actually increase for control speakers. Loss of "gay voice" has implications for accessibility. We also find that for speakers with "gay voice", loss of "gay voice" corresponds to lower similarity ratings. However, we caution that improving the ability of such models to synthesise ``gay voice'' comes with a great number of risks. We use this pipeline as a starting point for a discussion on the ethics of modelling queer voices more broadly. Collecting "clean" queer data has safety and fairness ramifications, and the resulting technology may cause harms from mockery to death.
Code (0)
등록된 구현이 없습니다.
Tasks
EthicsFairnessVoice CloningSimilar Papers 제목 키워드 기반
Long Scale Error Control in Low Light Image and Video Enhancement Using Equivariance
Image frames obtained in darkness are special. Just multiplying by a constant doesn't restore the image. Shot noise, quantization effects and camera non-linearities mean that colors and relative light levels are estimate…
QuantizationVideo EnhancementVideo Restoration`Just because you are right, doesn't mean I am wrong': Overcoming a bottleneck in development and evaluation of Open-Ended VQA tasks
GQA (CITATION) is a dataset for real-world visual reasoning and compositional question answering. We found that many answers predicted by the best vision-language models on the GQA dataset do not match the ground-truth a…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual ReasoningThe Effects of Just-in-time Delivery on Social Engagement: A Cluster Analysis
Fooji Inc. is a social media engagement platform that has created a proprietary "Just-in-time" delivery network to provide prizes to social media marketing campaign participants in real-time. In this paper, we prove the …
ClusteringMarketing'Just because you are right, doesn't mean I am wrong': Overcoming a Bottleneck in the Development and Evaluation of Open-Ended Visual Question Answering (VQA) Tasks
GQA~\citep{hudson2019gqa} is a dataset for real-world visual reasoning and compositional question answering. We found that many answers predicted by the best vision-language models on the GQA dataset do not match the gro…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual ReasoningWeakly Supervised Faster-RCNN+FPN to classify animals in camera trap images
Camera traps have revolutionized the animal research of many species that were previously nearly impossible to observe due to their habitat or behavior. They are cameras generally fixed to a tree that take a short sequen…
image-classificationImage Classificationobject-detectionObject Detection