paper-with-me

홈 › Papers

PWESuite: Phonetic Word Embeddings and Tasks They Facilitate

2023-04-05 · Vilém Zouhar, Kalvin Chang, Chenxuan Cui, Nathaniel Carlson, Nathaniel Robinson, Mrinmaya Sachan, David Mortensen

Mapping words into a fixed-dimensional vector space is the backbone of modern NLP. While most word embedding methods successfully encode semantic information, they overlook phonetic information that is crucial for many tasks. We develop three methods that use articulatory features to build phonetically informed word embeddings. To address the inconsistent evaluation of existing phonetic word embedding methods, we also contribute a task suite to fairly evaluate past, current, and future methods. We evaluate both (1) intrinsic aspects of phonetic word embeddings, such as word retrieval and correlation with sound similarity, and (2) extrinsic performance on tasks such as rhyme and cognate detection and sound analogies. We hope our task suite will promote reproducibility and inspire future phonetic embedding research.

📄 PDF Abstract BibTeX arXiv:2304.02541

Code (1)

zouharvi/pwesuite 공식 구현 pytorch

Tasks

RetrievalWord Embeddings

Similar Papers 제목 키워드 기반

Phonetic-and-Semantic Embedding of Spoken Words with Applications in Spoken Content Retrieval

2018-07-21 · Yi-Chen Chen, Sung-Feng Huang, Chia-Hao Shen, Hung-Yi Lee 외

Word embedding or Word2Vec has been successful in offering semantics for text words learned from the context of words. Audio Word2Vec was shown to offer phonetic structures for spoken words (signal segments for words) le…

Retrieval

Multilingual Jointly Trained Acoustic and Written Word Embeddings

2020-06-24 · Yushi Hu, Shane Settle, Karen Livescu

Acoustic word embeddings (AWEs) are vector representations of spoken word segments. AWEs can be learned jointly with embeddings of character sequences, to generate phonetically meaningful embeddings of written words, or …

Dynamic Time WarpingRetrievalWord Embeddings

Almost-unsupervised Speech Recognition with Close-to-zero Resource Based on Phonetic Structures Learned from Very Small Unpaired Speech and Text Data

2018-10-30 · Yi-Chen Chen, Chia-Hao Shen, Sung-Feng Huang, Hung-Yi Lee 외

Producing a large amount of annotated speech data for training ASR systems remains difficult for more than 95% of languages all over the world which are low-resourced. However, we note human babies start to learn the lan…

speech-recognitionSpeech RecognitionUnsupervised Speech Recognition

Leveraging multilingual transfer for unsupervised semantic acoustic word embeddings

2023-07-05 · Christiaan Jacobs, Herman Kamper

Acoustic word embeddings (AWEs) are fixed-dimensional vector representations of speech segments that encode phonetic content so that different realisations of the same word have similar embeddings. In this paper we explo…

Word EmbeddingsWord Similarity

Spoken Word2Vec: Learning Skipgram Embeddings from Speech

2023-11-15 · Mohammad Amaan Sayeed, Hanan Aldarmaki

Text word embeddings that encode distributional semantics work by modeling contextual similarities of frequently occurring words. Acoustic word embeddings, on the other hand, typically encode low-level phonetic similarit…

ClusteringWord Embeddings