paper-with-me

홈 › Papers

An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification

2023-05-22 · Yafeng Chen, Siqi Zheng, Hui Wang, Luyao Cheng, Qian Chen, Jiajun Qi

Effective fusion of multi-scale features is crucial for improving speaker verification performance. While most existing methods aggregate multi-scale features in a layer-wise manner via simple operations, such as summation or concatenation. This paper proposes a novel architecture called Enhanced Res2Net (ERes2Net), which incorporates both local and global feature fusion techniques to improve the performance. The local feature fusion (LFF) fuses the features within one single residual block to extract the local signal. The global feature fusion (GFF) takes acoustic features of different scales as input to aggregate global signal. To facilitate effective feature fusion in both LFF and GFF, an attentional feature fusion module is employed in the ERes2Net architecture, replacing summation or concatenation operations. A range of experiments conducted on the VoxCeleb datasets demonstrate the superiority of the ERes2Net in speaker verification. Code has been made publicly available at https://github.com/alibaba-damo-academy/3D-Speaker.

📄 PDF Abstract BibTeX arXiv:2305.12838

Code (2)

alibaba-damo-academy/3D-Speaker 공식 구현 pytorch
MindSpore-paper-code-3/code5/tree/main/res2net mindspore

Tasks

Speaker Verification

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Average Pooling 설명 없음
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Global Average Pooling Global Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding…
Kaiming Initialization 설명 없음
Res2Net Block A Res2Net Block is an image model block that constructs hierarchical residual-like connections within one single [residual…
Res2Net Res2Net is an image model that employs a variation on bottleneck residual blocks. The motivation is to be able to represent features at multiple scales. This is achieved…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency

2024-06-04 · Yafeng Chen, Siqi Zheng, Hui Wang, Luyao Cheng 외

Speaker verification systems experience significant performance degradation when tasked with short-duration trial recordings. To address this challenge, a multi-scale feature fusion approach has been proposed to effectiv…

Computational EfficiencySpeaker Verification

Global-Local Distillation Network-Based Audio-Visual Speaker Tracking with Incomplete Modalities

2024-08-26 · Yidi Li, Yihan Li, Yixin Guo, Bin Ren 외

In speaker tracking research, integrating and complementing multi-modal data is a crucial strategy for improving the accuracy and robustness of tracking systems. However, tracking with incomplete modalities remains a cha…

Generative Adversarial Network

STNet: Deep Audio-Visual Fusion Network for Robust Speaker Tracking

2024-10-08 · Yidi Li, Hong Liu, Bing Yang

Audio-visual speaker tracking aims to determine the location of human targets in a scene using signals captured by a multi-sensor platform, whose accuracy and robustness can be improved by multi-modal fusion methods. Rec…

Improving Transformer-based Networks With Locality For Automatic Speaker Verification

2023-02-17 · Mufan Sang, Yong Zhao, Gang Liu, John H. L. Hansen 외

Recently, Transformer-based architectures have been explored for speaker embedding extraction. Although the Transformer employs the self-attention mechanism to efficiently model the global interaction between token embed…

Speaker Verification

Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction

2023-12-16 · Zhaoxi Mu, Xinyu Yang, Sining Sun, Qing Yang

Speech signals are inherently complex as they encompass both global acoustic characteristics and local semantic information. However, in the task of target speech extraction, certain elements of global and local semantic…

DisentanglementRepresentation LearningSpeech Extraction