Unsupervised Segmentation of Phoneme Sequences based on Pitman-Yor Semi-Markov Model using Phoneme Length Context
Unsupervised segmentation of phoneme sequences is an essential process to obtain unknown words during spoken dialogues. In this segmentation, an input phoneme sequence without delimiters is converted into segmented sub-sequences corresponding to words. The Pitman-Yor semi-Markov model (PYSMM) is promising for this problem, but its performance degrades when it is applied to phoneme-level word segmentation. This is because of insufficient cues for the segmentation, e.g., homophones are improperly treated as single entries and their different contexts are also confused. We propose a phoneme-length context model for PYSMM to give a helpful cue at the phoneme-level and to predict succeeding segments more accurately. Our experiments showed that the peak performance with our context model outperformed those without such a context model by 0.045 at most in terms of F-measures of estimated segmentation.
Code (0)
등록된 구현이 없습니다.
Tasks
SegmentationSpoken Dialogue SystemsSimilar Papers 제목 키워드 기반
Une m\'ethode non-supervis\'ee pour la segmentation morphologique et l'apprentissage de morphotactique \`a l'aide de processus de Pitman-Yor (An unsupervised method for joint morphological segmentation and morphotactics learning using Pitman-Yor processes)
Cet article pr{\'e}sente un mod{\`e}le bay{\'e}sien non-param{\'e}trique pour la segmentation morphologique non supervis{\'e}e. Ce mod{\`e}le semi-markovien s{'}appuie sur des classes latentes de morph{\`e}mes afin de mo…
MORPHUnsupervised Melody Segmentation Based on a Nested Pitman-Yor Language Model
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
We introduce DiscoPhon, a multilingual benchmark for evaluating unsupervised phoneme discovery from discrete speech units. DiscoPhon covers 6 dev and 6 test languages, chosen to span a wide range of phonemic contrasts. G…
Completely Unsupervised Phoneme Recognition by Adversarially Learning Mapping Relationships from Audio Embeddings
Unsupervised discovery of acoustic tokens from audio corpora without annotation and learning vector representations for these tokens have been widely studied. Although these techniques have been shown successful in some …
Generative Adversarial NetworkPhoneme RecognitionTowards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning
In this paper we propose a Sequential Representation Quantization AutoEncoder (SeqRQ-AE) to learn from primarily unpaired audio data and produce sequences of representations very close to phoneme sequences of speech utte…
ClusteringPhoneme RecognitionQuantizationRepresentation Learning+4