paper-with-me

홈 › Papers

Contextual Joint Factor Acoustic Embeddings

2019-10-16 · Yanpei Shi, Thomas Hain

Embedding acoustic information into fixed length representations is of interest for a whole range of applications in speech and audio technology. Two novel unsupervised approaches to generate acoustic embeddings by modelling of acoustic context are proposed. The first approach is a contextual joint factor synthesis encoder, where the encoder in an encoder/decoder framework is trained to extract joint factors from surrounding audio frames to best generate the target output. The second approach is a contextual joint factor analysis encoder, where the encoder is trained to analyse joint factors from the source signal that correlates best with the neighbouring audio. To evaluate the effectiveness of our approaches compared to prior work, two tasks are conducted -- phone classification and speaker recognition -- and test on different TIMIT data sets. Experimental results show that one of the proposed approaches outperforms phone classification baselines, yielding a classification accuracy of 74.1%. When using additional out-of-domain data for training, an additional 3% improvements can be obtained, for both for phone classification and speaker recognition tasks.

📄 PDF Abstract BibTeX arXiv:1910.07601

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationDecoderGeneral ClassificationSpeaker Recognition

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

attr2vec: Jointly Learning Word and Contextual Attribute Embeddings with Factorization Machines

2018-06-01 · NAACL 2018 6 · Fabio Petroni, Vassilis Plachouras, Timothy Nugent, Jochen L. Leidner

The widespread use of word embeddings is associated with the recent successes of many natural language processing (NLP) systems. The key approach of popular models such as word2vec and GloVe is to learn dense vector repr…

AttributeDependency ParsingInformation RetrievalMachine Translation+5

Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach

2023-09-11 · Tae Jin Park, Kunal Dhawan, Nithin Koluguri, Jagadeesh Balam

Large language models (LLMs) have shown great promise for capturing contextual information in natural language processing tasks. We propose a novel approach to speaker diarization that incorporates the prowess of LLMs to…

speaker-diarizationSpeaker Diarization

Learned in Speech Recognition: Contextual Acoustic Word Embeddings

2018-10-22 · Anonymous

End-to-end acoustic-to-word speech recognition models have recently gained popularity because they are easy to train, scale well to large amounts of training data, and do not require a lexicon. In addition, word models m…

Sentencespeech-recognitionSpeech RecognitionSpoken Language Understanding+1

Learned In Speech Recognition: Contextual Acoustic Word Embeddings

2019-02-18 · Shruti Palaskar, Vikas Raunak, Florian Metze

End-to-end acoustic-to-word speech recognition models have recently gained popularity because they are easy to train, scale well to large amounts of training data, and do not require a lexicon. In addition, word models m…

Sentencespeech-recognitionSpeech RecognitionSpoken Language Understanding+1

CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR

2026-01-30 · Muhammad Shakeel, Yosuke Fukumoto, Chikara Maeda, Chyi-Jiunn Lin 외 arxiv

We present CALM, a joint Contextual Acoustic-Linguistic Modeling framework for multi-speaker automatic speech recognition (ASR). In personalized AI scenarios, the joint availability of acoustic and linguistic cues natura…

Speech Recognition