paper-with-me

홈 › Papers

Timestamped Embedding-Matching Acoustic-to-Word CTC ASR

2023-06-20 · Woojay Jeon

In this work, we describe a novel method of training an embedding-matching word-level connectionist temporal classification (CTC) automatic speech recognizer (ASR) such that it directly produces word start times and durations, required by many real-world applications, in addition to the transcription. The word timestamps enable the ASR to output word segmentations and word confusion networks without relying on a secondary model or forced alignment process when testing. Our proposed system has similar word segmentation accuracy as a hybrid DNN-HMM (Deep Neural Network-Hidden Markov Model) system, with less than 3ms difference in mean absolute error in word start times on TIMIT data. At the same time, we observed less than 5% relative increase in the word error rate compared to the non-timestamped system when using the same audio training data and nearly identical model size. We also contribute more rigorous analysis of multiple-hypothesis embedding-matching ASR in general.

📄 PDF Abstract BibTeX arXiv:2306.11473

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Visualization of Change in Word Meaning over Time using Temporal Word Embeddings

2014-10-18 · Chiraag Lala, Shay B. Cohen

We describe a visualization tool that can be used to view the change in meaning of words over time. The tool makes use of existing (static) word embedding datasets together with a timestamped $n$-gram corpus to create {\…

Word Embeddings

Query-by-Example Search with Discriminative Neural Acoustic Word Embeddings

2017-06-12 · Shane Settle, Keith Levin, Herman Kamper, Karen Livescu

Query-by-example search often uses dynamic time warping (DTW) for comparing queries and proposed matching segments. Recent work has shown that comparing speech segments by representing them as fixed-dimensional vectors -…

Dynamic Time WarpingWord Embeddings

Learning Acoustic Word Embeddings with Temporal Context for Query-by-Example Speech Search

2018-06-10 · Yougen Yuan, Cheung-Chi Leung, Lei Xie, Hongjie Chen 외

We propose to learn acoustic word embeddings with temporal context for query-by-example (QbE) speech search. The temporal context includes the leading and trailing word sequences of a word. We assume that there exist spo…

Dynamic Time WarpingTripletWord Embeddings

Improvements to Embedding-Matching Acoustic-to-Word ASR Using Multiple-Hypothesis Pronunciation-Based Embeddings

2022-10-30 · Hao Yen, Woojay Jeon

In embedding-matching acoustic-to-word (A2W) ASR, every word in the vocabulary is represented by a fixed-dimension embedding vector that can be added or removed independently of the rest of the system. The approach is po…

Acoustic Neighbor Embeddings

2020-07-20 · Woojay Jeon

This paper proposes a novel acoustic word embedding called Acoustic Neighbor Embeddings where speech or text of arbitrary length are mapped to a vector space of fixed, reduced dimensions by adapting stochastic neighbor e…

Triplet