Are we describing the same sound? An analysis of word embedding spaces of expressive piano performance
Semantic embeddings play a crucial role in natural language-based information retrieval. Embedding models represent words and contexts as vectors whose spatial configuration is derived from the distribution of words in large text corpora. While such representations are generally very powerful, they might fail to account for fine-grained domain-specific nuances. In this article, we investigate this uncertainty for the domain of characterizations of expressive piano performance. Using a music research dataset of free text performance characterizations and a follow-up study sorting the annotations into clusters, we derive a ground truth for a domain-specific semantic similarity structure. We test five embedding models and their similarity structure for correspondence with the ground truth. We further assess the effects of contextualizing prompts, hubness reduction, cross-modal similarity, and k-means clustering. The quality of embedding models shows great variability with respect to this task; more general models perform better than domain-adapted ones and the best model configurations reach human-level agreement.
Code (1)
Tasks
Information RetrievalRetrievalSemantic SimilaritySemantic Textual SimilaritySimilar Papers 제목 키워드 기반
Modeling Personal Biases in Language Use by Inducing Personalized Word Embeddings
There exist biases in individual{'}s language use; the same word (e.g., cool) is used for expressing different meanings (e.g., temperature range) or different words (e.g., cloudy, hazy) are used for describing the same m…
Multi-class ClassificationMulti-Task LearningSentiment AnalysisWord EmbeddingsText-Driven Separation of Arbitrary Sounds
We propose a method of separating a desired sound source from a single-channel mixture, based on either a textual description or a short audio sample of the target source. This is achieved by combining two distinct model…
Learning Topic Models by Neighborhood Aggregation
Topic models are frequently used in machine learning owing to their high interpretability and modular structure. However, extending a topic model to include a supervisory signal, to incorporate pre-trained word embedding…
Document ClassificationTopic ModelsIndian EmoSpeech Command Dataset: A dataset for emotion based speech recognition in the wild
Speech emotion analysis is an important task which further enables several application use cases. The non-verbal sounds within speech utterances also play a pivotal role in emotion analysis in speech. Due to the widespre…
Emotion RecognitionKeyword Spottingspeech-recognitionSpeech RecognitionSound-Word2Vec: Learning Word Representations Grounded in Sounds
To be able to interact better with humans, it is crucial for machines to understand sound - a primary modality of human perception. Previous works have used sound to learn embeddings for improved generic textual similari…
RetrievalWord Embeddings