paper-with-me

Papers

Metric Learning for Phoneme Perception

2018-09-20 · Yair Lakretz, Gal Chechik, Evan-Gary Cohen, Alessandro Treves, Naama Friedmann

Metric functions for phoneme perception capture the similarity structure among phonemes in a given language and therefore play a central role in phonology and psycho-linguistics. Various phenomena depend on phoneme similarity, such as spoken word recognition or serial recall from verbal working memory. This study presents a new framework for learning a metric function for perceptual distances among pairs of phonemes. Previous studies have proposed various metric functions, from simple measures counting the number of phonetic dimensions that two phonemes share (place-, manner-of-articulation and voicing), to more sophisticated ones such as deriving perceptual distances based on the number of natural classes that both phonemes belong to. However, previous studies have manually constructed the metric function, which may lead to unsatisfactory account of the empirical data. This study presents a framework to derive the metric function from behavioral data on phoneme perception using learning algorithms. We first show that this approach outperforms previous metrics suggested in the literature in predicting perceptual distances among phoneme pairs. We then study several metric functions derived by the learning algorithms and show how perceptual saliencies of phonological features can be derived from them. For English, we show that the derived perceptual saliencies are in accordance with a previously described order among phonological features and show how the framework extends the results to more features. Finally, we explore how the metric function and perceptual saliencies of phonological features may vary across languages. To this end, we compare results based on two English datasets and a new dataset that we have collected for Hebrew.

📄 PDF Abstract BibTeX arXiv:1809.07824

Code (0)

등록된 구현이 없습니다.

Tasks

Metric Learning

Similar Papers 제목 키워드 기반

The Lipreading Gap: Do VSR Models Perceive Visual Speech Like Human Lipreaders?

2026-06-05 · Rishabh Jain, Naomi Harte arxiv

Visual speech recognition (VSR) models now surpass human lipreaders on benchmarks, but do such gains establish human-like visual speech perception? To explore this, we compare three VSR systems with human baselines on th…

Visual Speech Recognition

Predicting non-native speech perception using the Perceptual Assimilation Model and state-of-the-art acoustic models

2022-05-31 · CoNLL (EMNLP) 2021 11 · Juliette Millet, Ioana Chitoran, Ewan Dunbar

Our native language influences the way we perceive speech sounds, affecting our ability to discriminate non-native sounds. We compare two ideas about the influence of the native language on speech perception: the Percept…

MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification

2026-02-27 · Liang Jinghua, Zhang Zifeng, Li Songyi, Zheng Linze arxiv

We propose MEBM-Phoneme, a multi-scale enhanced neural decoder for phoneme classification from non-invasive magnetoencephalography (MEG) signals. Built upon the BrainMagic backbone, MEBM-Phoneme integrates a short-term m…

Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction

2025-01-08 · JIhwan Lee, Tiantian Feng, Aditya Kommineni, Sudarsana Reddy Kadiri 외

Brain-computer interfaces (BCI) offer numerous human-centered application possibilities, particularly affecting people with neurological disorders. Text or speech decoding from brain activities is a relevant domain that …

EEG

Estimating Phoneme Class Conditional Probabilities from Raw Speech Signal using Convolutional Neural Networks

2013-04-03 · Dimitri Palaz, Ronan Collobert, Mathew Magimai. -Doss

In hybrid hidden Markov model/artificial neural networks (HMM/ANN) automatic speech recognition (ASR) system, the phoneme class conditional probabilities are estimated by first extracting acoustic features from the speec…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Phoneme Recognitionspeech-recognition+1