Familiar words but strange voices: Modelling the influence of speech variability on word recognition
We present a deep neural model of spoken word recognition which is trained to retrieve the meaning of a word (in the form of a word embedding) given its spoken form, a task which resembles that faced by a human listener. Furthermore, we investigate the influence of variability in speech signals on the model{'}s performance. To this end, we conduct of set of controlled experiments using word-aligned read speech data in German. Our experiments show that (1) the model is more sensitive to dialectical variation than gender variation, and (2) recognition performance of word cognates from related languages reflect the degree of relatedness between languages in our study. Our work highlights the feasibility of modeling human speech perception using deep neural networks.
Code (0)
등록된 구현이 없습니다.
Tasks
FormSimilar Papers 제목 키워드 기반
The Impact of a Chatbot's Ephemerality-Framing on Self-Disclosure Perceptions
Self-disclosure, the sharing of one's thoughts and feelings, is affected by the perceived relationship between individuals. While chatbots are increasingly used for self-disclosure, the impact of a chatbot's framing on u…
ChatbotStrangeness-driven Exploration in Multi-Agent Reinforcement Learning
Efficient exploration strategy is one of essential issues in cooperative multi-agent reinforcement learning (MARL) algorithms requiring complex coordination. In this study, we introduce a new exploration method with the …
Efficient ExplorationMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning+2Artificial Intelligence as Strange Intelligence: Against Linear Models of Intelligence
We endorse and expand upon Susan Schneider's critique of the linear model of AI progress and introduce two novel concepts: "familiar intelligence" and "strange intelligence". AI intelligence is likely to be strange intel…
An overview of text-to-speech systems and media applications
Producing synthetic voice, similar to human-like sound, is an emerging novelty of modern interactive media systems. Text-To-Speech (TTS) systems try to generate synthetic and authentic voices via text input. Besides, wel…
Acoustic Modellingtext-to-speechText to SpeechVoice ConversionEstranged Predictions: Measuring Semantic Category Disruption with Masked Language Modelling
This paper examines how science fiction destabilises ontological categories by measuring conceptual permeability across the terms human, animal, and machine using masked language modelling (MLM). Drawing on corpora of sc…
Language Modelling