paper-with-me

홈 › Papers

Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks

2023-09-14 · Sizhou Chen, Songyang Gao, Sen Fang

The Transformer architecture has proven to be highly effective for Automatic Speech Recognition (ASR) tasks, becoming a foundational component for a plethora of research in the domain. Historically, many approaches have leaned on fixed-length attention windows, which becomes problematic for varied speech samples in duration and complexity, leading to data over-smoothing and neglect of essential long-term connectivity. Addressing this limitation, we introduce Echo-MSA, a nimble module equipped with a variable-length attention mechanism that accommodates a range of speech sample complexities and durations. This module offers the flexibility to extract speech features across various granularities, spanning from frames and phonemes to words and discourse. The proposed design captures the variable length feature of speech and addresses the limitations of fixed-length attention. Our evaluation leverages a parallel attention architecture complemented by a dynamic gating mechanism that amalgamates traditional attention with the Echo-MSA module output. Empirical evidence from our study reveals that integrating Echo-MSA into the primary model's training regime significantly enhances the word error rate (WER) performance, all while preserving the intrinsic stability of the original model.

📄 PDF Abstract BibTeX arXiv:2309.07765

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Adam 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

ACA-Net: Towards Lightweight Speaker Verification using Asymmetric Cross Attention

2023-05-20 · Jia Qi Yip, Tuan Truong, Dianwen Ng, Chong Zhang 외

In this paper, we propose ACA-Net, a lightweight, global context-aware speaker embedding extractor for Speaker Verification (SV) that improves upon existing work by using Asymmetric Cross Attention (ACA) to replace tempo…

Speaker Verification

REFACTOR: Learning to Extract Theorems from Proofs

2024-02-26 · Jin Peng Zhou, Yuhuai Wu, Qiyang Li, Roger Grosse

Human mathematicians are often good at recognizing modular and reusable theorems that make complex mathematical results within reach. In this paper, we propose a novel method called theoREm-from-prooF extrACTOR (REFACTOR…

Automated Theorem Proving

Speaker Characterization by means of Attention Pooling

2024-05-07 · Federico Costa, Miquel India, Javier Hernando

State-of-the-art Deep Learning systems for speaker verification are commonly based on speaker embedding extractors. These architectures are usually composed of a feature extractor front-end together with a pooling layer …

Emotion RecognitionSpeaker RecognitionSpeaker Verification

SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning

2021-11-25 · CVPR 2022 1 · Kevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed 외

The canonical approach to video captioning dictates a caption generation model to learn from offline-extracted dense video features. These feature extractors usually operate on video frames sampled at a fixed frame rate …

Caption GenerationQuestion AnsweringVideo CaptioningVideo Question Answering+1

Attention Is Not All You Need Anymore

2023-08-15 · Zhe Chen

In recent years, the popular Transformer architecture has achieved great success in many application areas, including natural language processing and computer vision. Many existing works aim to reduce the computational a…

AllText Generation