paper-with-me

홈 › Papers

Neural Network based End-to-End Query by Example Spoken Term Detection

2019-11-19 · Dhananjay Ram, Lesly Miculicich, Hervé Bourlard

This paper focuses on the problem of query by example spoken term detection (QbE-STD) in zero-resource scenario. State-of-the-art approaches primarily rely on dynamic time warping (DTW) based template matching techniques using phone posterior or bottleneck features extracted from a deep neural network (DNN). We use both monolingual and multilingual bottleneck features, and show that multilingual features perform increasingly better with more training languages. Previously, it has been shown that the DTW based matching can be replaced with a CNN based matching while using posterior features. Here, we show that the CNN based matching outperforms DTW based matching using bottleneck features as well. In this case, the feature extraction and pattern matching stages of our QbE-STD system are optimized independently of each other. We propose to integrate these two stages in a fully neural network based end-to-end learning framework to enable joint optimization of those two stages simultaneously. The proposed approaches are evaluated on two challenging multilingual datasets: Spoken Web Search 2013 and Query by Example Search on Speech Task 2014, demonstrating in each case significant improvements.

📄 PDF Abstract BibTeX arXiv:1911.08332

Code (0)

등록된 구현이 없습니다.

Tasks

Dynamic Time WarpingTemplate Matching

Methods 이 논문이 사용한 방법론

DTW Dynamic Time Warping (DTW) [1] is one of well-known distance measures between a pairwise of time series. The main idea of DTW is to compute the distance from the matching of…

Similar Papers 제목 키워드 기반

H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing

2025-06-20 · Akanksha Singh, Yi-Ping Phoebe Chen, Vipul Arora

Query-by-example spoken term detection (QbE-STD) searches for matching words or phrases in an audio dataset using a sample spoken query. When annotated data is limited or unavailable, QbE-STD is often done using template…

Dynamic Time WarpingRepresentation LearningRetrievalTemplate Matching

Query-by-example Spoken Term Detection using Attention-based Multi-hop Networks

2017-09-01 · Chia-Wei Ao, Hung-Yi Lee

Retrieving spoken content with spoken queries, or query-by- example spoken term detection (STD), is attractive because it makes possible the matching of signals directly on the acoustic level without transcribing them in…

NTU System at MediaEval 2015: Zero Resource Query by Example Spoken Term Detection Using Deep and Recurrent Neural Networks

2015-09-14 · MediaEval 2015 Workshop 2015 9 · Cheng-Tao Chung, Yang-De Chen

This note serves as a documentation describing the methods the authors of this paper implemented for the Query by Example Search on Speech Task (QUESST) as a part of MediaEval 2015. In this work, we combined DTW, DNN and…

Keyword Spotting

Cross-Lingual Query-by-Example Spoken Term Detection: A Transformer-Based Approach

2024-10-05 · Allahdadi Fatemeh, Mahdian Toroghi Rahil, Zareian Hassan

Query-by-example spoken term detection (QbE-STD) is typically constrained by transcribed data scarcity and language specificity. This paper introduces a novel, language-agnostic QbE-STD model leveraging image processing …

Specificity

Use of GPU and Feature Reduction for Fast Query-by-Example Spoken Term Detection

2014-12-01 · WS 2014 12 · Gautam Mantena, Kishore Prahallad
GPUSpeech Recognition