paper-with-me

Papers

An Acoustic Segment Model Based Segment Unit Selection Approach to Acoustic Scene Classification with Partial Utterances

2020-07-31 · Hu Hu, Sabato Marco Siniscalchi, Yannan Wang, Xue Bai, Jun Du, Chin-Hui Lee

In this paper, we propose a sub-utterance unit selection framework to remove acoustic segments in audio recordings that carry little information for acoustic scene classification (ASC). Our approach is built upon a universal set of acoustic segment units covering the overall acoustic scene space. First, those units are modeled with acoustic segment models (ASMs) used to tokenize acoustic scene utterances into sequences of acoustic segment units. Next, paralleling the idea of stop words in information retrieval, stop ASMs are automatically detected. Finally, acoustic segments associated with the stop ASMs are blocked, because of their low indexing power in retrieval of most acoustic scenes. In contrast to building scene models with whole utterances, the ASM-removed sub-utterances, i.e., acoustic utterances without stop acoustic segments, are then used as inputs to the AlexNet-L back-end for final classification. On the DCASE 2018 dataset, scene classification accuracy increases from 68%, with whole utterances, to 72.1%, with segment selection. This represents a competitive accuracy without any data augmentation, and/or ensemble strategy. Moreover, our approach compares favourably to AlexNet-L with attention.

📄 PDF Abstract BibTeX arXiv:2008.00107

Code (0)

등록된 구현이 없습니다.

Tasks

Acoustic Scene ClassificationClassificationData AugmentationGeneral ClassificationInformation RetrievalRetrievalScene Classification

Similar Papers 제목 키워드 기반

Segmenting Subtitles for Correcting ASR Segmentation Errors

2021-04-16 · EACL 2021 2 · David Wan, Chris Kedzie, Faisal Ladhak, Elsbeth Turcan 외

Typical ASR systems segment the input audio into utterances using purely acoustic information, which may not resemble the sentence-like units that are expected by conventional machine translation (MT) systems for Spoken …

Information RetrievalMachine TranslationRetrievalSegmentation+2

Whole-Word Segmental Speech Recognition with Acoustic Word Embeddings

2020-07-01 · Bowen Shi, Shane Settle, Karen Livescu

Segmental models are sequence prediction models in which scores of hypotheses are based on entire variable-length segments of frames. We consider segmental models for whole-word ("acoustic-to-word") speech recognition, w…

GPUspeech-recognitionSpeech RecognitionWord Embeddings

Acoustic Data-Driven Subword Modeling for End-to-End Speech Recognition

2021-04-19 · Wei Zhou, Mohammad Zeineldeen, Zuoyun Zheng, Ralf Schlüter 외

Subword units are commonly used for end-to-end automatic speech recognition (ASR), while a fully acoustic-oriented subword modeling approach is somewhat missing. We propose an acoustic data-driven subword modeling (ADSM)…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Segmentationspeech-recognition+2

Statistical Context-Dependent Units Boundary Correction for Corpus-based Unit-Selection Text-to-Speech

2020-03-05 · Claudio Zito, Fabio Tesser, Mauro Nicolao, Piero Cosi

In this study, we present an innovative technique for speaker adaptation in order to improve the accuracy of segmentation with application to unit-selection Text-To-Speech (TTS) systems. Unlike conventional techniques fo…

Segmentationtext-to-speechText to Speech

Revisiting speech segmentation and lexicon learning with better features

2024-01-31 · Herman Kamper, Benjamin van Niekerk

We revisit a self-supervised method that segments unlabelled speech into word-like segments. We start from the two-stage duration-penalised dynamic programming method that performs zero-resource segmentation without lear…

Acoustic Unit DiscoverySegmentation