paper-with-me

Papers

Frequency and temporal convolutional attention for text-independent speaker recognition

2019-10-16 · Sarthak Yadav, Atul Rai

Majority of the recent approaches for text-independent speaker recognition apply attention or similar techniques for aggregation of frame-level feature descriptors generated by a deep neural network (DNN) front-end. In this paper, we propose methods of convolutional attention for independently modelling temporal and frequency information in a convolutional neural network (CNN) based front-end. Our system utilizes convolutional block attention modules (CBAMs) [1] appropriately modified to accommodate spectrogram inputs. The proposed CNN front-end fitted with the proposed convolutional attention modules outperform the no-attention and spatial-CBAM baselines by a significant margin on the VoxCeleb [2, 3] speaker verification benchmark, and our best model achieves an equal error rate of 2:031% on the VoxCeleb1 test set, improving the existing state of the art result by a significant margin. For a more thorough assessment of the effects of frequency and temporal attention in real-world conditions, we conduct ablation experiments by randomly dropping frequency bins and temporal frames from the input spectrograms, concluding that instead of modelling either of the entities, simultaneously modelling temporal and frequency attention translates to better real-world performance.

📄 PDF Abstract BibTeX arXiv:1910.07364

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker RecognitionSpeaker VerificationText-Independent Speaker Recognition

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Prosodic-Enhanced Siamese Convolutional Neural Networks for Cross-Device Text-Independent Speaker Verification

2018-07-31 · Sobhan Soleymani, Ali Dabouei, Seyed Mehdi Iranmanesh, Hadi Kazemi 외

In this paper a novel cross-device text-independent speaker verification architecture is proposed. Majority of the state-of-the-art deep architectures that are used for speaker verification tasks consider Mel-frequency c…

Speaker VerificationText-Independent Speaker Verification

MFA: TDNN with Multi-scale Frequency-channel Attention for Text-independent Speaker Verification with Short Utterances

2022-02-03 · Tianchi Liu, Rohan Kumar Das, Kong Aik Lee, Haizhou Li

The time delay neural network (TDNN) represents one of the state-of-the-art of neural solutions to text-independent speaker verification. However, they require a large number of filters to capture the speaker characteris…

Speaker VerificationText-Independent Speaker Verification

I\textsuperscript{2}RiMA: Spectral Riemannian Representation with Temporal Attention for Mental Stress Detection based on EEG Signals

2026-07-01 · Cheng He, Kunyu Peng, Shangen Han, Jinming Ma 외 arxiv

Cross-subject EEG stress detection remains challenging because discriminative stress-related patterns are both subject-dependent and frequency-specific. Conventional Riemannian methods model spatial covariance mainly in …

Divided spectro-temporal attention for sound event localization and detection in real scenes for DCASE2023 challenge

2023-06-05 · Yusun Shul, Byeong-Yun Ko, Jung-Woo Choi

Localizing sounds and detecting events in different room environments is a difficult task, mainly due to the wide range of reflections and reverberations. When training neural network models with sounds recorded in only …

Event DetectionSound Event DetectionSound Event Localization and Detection

ACFormer: Mitigating Non-linearity with Auto Convolutional Encoder for Time Series Forecasting

2026-01-28 · Gawon Lee, Hanbyeol Park, Minseop Kim, Dohee Kim 외 arxiv

Time series forecasting (TSF) faces challenges in modeling complex intra-channel temporal dependencies and inter-channel correlations. Although recent research has highlighted the efficiency of linear architectures in ca…

Time Series Forecasting