paper-with-me

Papers

Multimodal and Multiresolution Speech Recognition with Transformers

2020-07-01 · ACL 2020 6 · Georgios Paraskevopoulos, Srinivas Parthasarathy, Aparna Khare, Shiva Sundaram

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract representations for audio features in the encoder layers of the transformer and fuse video features using an additional crossmodal multihead attention layer. Additionally, we incorporate a multitask training criterion for multiresolution ASR, where we train the model to generate both character and subword level transcriptions. Experimental results on the How2 dataset, indicate that multiresolution training can speed up convergence by around 50{\%} and relatively improves word error rate (WER) performance by upto 18{\%} over subword prediction models. Further, incorporating visual information improves performance with relative gains upto 3.76{\%} over audio only models. Our results are comparable to state-of-the-art Listen, Attend and Spell-based architectures.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
SentencePiece 설명 없음
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

Multiresolution and Multimodal Speech Recognition with Transformers

2020-04-29 · Georgios Paraskevopoulos, Srinivas Parthasarathy, Aparna Khare, Shiva Sundaram

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. W…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Transformers in Speech Processing: A Survey

2023-03-21 · Siddique Latif, Aun Zaidi, Heriberto Cuayahuitl, Fahad Shamshad 외

The remarkable success of transformers in the field of natural language processing has sparked the interest of the speech-processing community, leading to an exploration of their potential for modeling long-range depende…

Automatic Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition+3

SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition

2025-02-01 · Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian Mishara

In the field of human-computer interaction and psychological assessment, speech emotion recognition (SER) plays an important role in deciphering emotional states from speech signals. Despite advancements, challenges pers…

DenoisingEmotion RecognitionSpeech Emotion Recognition

Multimodal Emotion Recognition with Transformer-Based Self Supervised Feature Fusion

2020-10-27 · Shamane Siriwardhana ; Tharindu Kaluarachchi ; Mark Billinghurst ; Suranga Nanayakkara

Emotion Recognition is a challenging research area given its complex nature, and humans express emotional cues across various modalities such as language, facial expressions, and speech. Representation and fusion of feat…

Emotion RecognitionMultimodal Deep LearningMultimodal Emotion RecognitionMultimodal Sentiment Analysis+2

Fast Boosting Based Detection Using Scale Invariant Multimodal Multiresolution Filtered Features

2017-07-01 · CVPR 2017 7 · Arthur Daniel Costea, Robert Varga, Sergiu Nedevschi

In this paper we propose a novel boosting-based sliding window solution for object detection which can keep up with the precision of the state-of-the art deep learning approaches, while being 10 to 100 times faster. The …

object-detectionObject Detection