paper-with-me

홈 › Papers

Leveraging Acoustic and Linguistic Embeddings from Pretrained speech and language Models for Intent Classification

2021-02-15 · Bidisha Sharma, Maulik Madhavi, Haizhou Li

Intent classification is a task in spoken language understanding. An intent classification system is usually implemented as a pipeline process, with a speech recognition module followed by text processing that classifies the intents. There are also studies of end-to-end system that takes acoustic features as input and classifies the intents directly. Such systems don't take advantage of relevant linguistic information, and suffer from limited training data. In this work, we propose a novel intent classification framework that employs acoustic features extracted from a pretrained speech recognition system and linguistic features learned from a pretrained language model. We use knowledge distillation technique to map the acoustic embeddings towards linguistic embeddings. We perform fusion of both acoustic and linguistic embeddings through cross-attention approach to classify intents. With the proposed method, we achieve 90.86% and 99.07% accuracy on ATIS and Fluent speech corpus, respectively.

📄 PDF Abstract BibTeX arXiv:2102.07370

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral Classificationintent-classificationIntent ClassificationKnowledge DistillationLanguage ModelingLanguage Modellingspeech-recognitionSpeech RecognitionSpoken Language Understanding

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Demographic Attributes Prediction from Speech Using WavLM Embeddings

2025-02-17 · Yuchen Yang, Thomas Thebaud, Najim Dehak

This paper introduces a general classifier based on WavLM features, to infer demographic characteristics, such as age, gender, native language, education, and country, from speech. Demographic feature prediction plays a …

DiversityGender ClassificationPrediction

Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection

2026-06-26 · Olivier Jiyoun Jung, Jonghyeon Park, Myungwoo Oh arxiv

Early detection of dementia through speech analysis offers a non-invasive screening alternative, but capturing both acoustic and linguistic biomarkers remains challenging. We propose a multimodal framework leveraging Whi…

Speech Recognition

Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue

2023-12-23 · Guan-Ting Lin, Prashanth Gurunath Shivakumar, Ankur Gandhe, Chao-Han Huck Yang 외

Large Language Models (LLMs) have demonstrated superior abilities in tasks such as chatting, reasoning, and question-answering. However, standard LLMs may ignore crucial paralinguistic information, such as sentiment, emo…

AttributeLanguage ModelingLanguage ModellingQuestion Answering+3

Audio-Linguistic Embeddings for Spoken Sentences

2019-02-20 · Albert Haque, Michelle Guo, Prateek Verma, Li Fei-Fei

We propose spoken sentence embeddings which capture both acoustic and linguistic content. While existing works operate at the character, phoneme, or word level, our method learns long-term dependencies by modeling speech…

DecoderEmotion RecognitionSentenceSentence Embeddings+3

CTA-RNN: Channel and Temporal-wise Attention RNN Leveraging Pre-trained ASR Embeddings for Speech Emotion Recognition

2022-03-31 · Chengxin Chen, Pengyuan Zhang

Previous research has looked into ways to improve speech emotion recognition (SER) by utilizing both acoustic and linguistic cues of speech. However, the potential association between state-of-the-art ASR models and the …

Cross-corpusEmotion RecognitionSpeech Emotion Recognition