paper-with-me

홈 › Papers

LaSAFT: Latent Source Attentive Frequency Transformation for Conditioned Source Separation

2020-10-22 · Woosung Choi, Minseok Kim, Jaehwa Chung, Soonyoung Jung

Recent deep-learning approaches have shown that Frequency Transformation (FT) blocks can significantly improve spectrogram-based single-source separation models by capturing frequency patterns. The goal of this paper is to extend the FT block to fit the multi-source task. We propose the Latent Source Attentive Frequency Transformation (LaSAFT) block to capture source-dependent frequency patterns. We also propose the Gated Point-wise Convolutional Modulation (GPoCM), an extension of Feature-wise Linear Modulation (FiLM), to modulate internal features. By employing these two novel methods, we extend the Conditioned-U-Net (CUNet) for multi-source separation, and the experimental results indicate that our LaSAFT and GPoCM can improve the CUNet's performance, achieving state-of-the-art SDR performance on several MUSDB18 source separation tasks.

📄 PDF Abstract BibTeX arXiv:2010.11631

Code (1)

ws-choi/Conditioned-Source-Separation-LaSAFT 공식 구현 pytorch

Tasks

Music Source Separation

Similar Papers 제목 키워드 기반

LightSAFT: Lightweight Latent Source Aware Frequency Transform for Source Separation

2021-11-24 · Yeong-Seok Jeong, Jinsung Kim, Woosung Choi, Jaehwa Chung 외

Conditioned source separations have attracted significant attention because of their flexibility, applicability and extensionality. Their performance was usually inferior to the existing approaches, such as the single so…

Attentive Normalization

2019-08-04 · ECCV 2020 8 · Xilai Li, Wei Sun, Tianfu Wu

In state-of-the-art deep neural networks, both feature normalization and feature attention have become ubiquitous. % with significant performance improvement shown in a vast amount of tasks. They are usually studied as s…

Image ClassificationInstance Segmentationobject-detectionObject Detection+1

HAZE-Net: High-Frequency Attentive Super-Resolved Gaze Estimation in Low-Resolution Face Images

2022-09-21 · Jun-Seok Yun, Youngju Na, Hee Hyeon Kim, Hyung-Il Kim 외

Although gaze estimation methods have been developed with deep learning techniques, there has been no such approach as aim to attain accurate performance in low-resolution face images with a pixel width of 50 pixels or l…

Gaze EstimationSuper-Resolution

Attentively Embracing Noise for Robust Latent Representation in BERT

2020-12-01 · COLING 2020 8 · Gwenaelle Cunha Sergio, Dennis Singh Moirangthem, Minho Lee

Modern digital personal assistants interact with users through voice. Therefore, they heavily rely on automatic speech recognition (ASR) in order to convert speech to text and perform further tasks. We introduce EBERT, w…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ChatbotClassification+7

Attentive Mimicking: Better Word Embeddings by Attending to Informative Contexts

2019-04-02 · NAACL 2019 6 · Timo Schick, Hinrich Schütze

Learning high-quality embeddings for rare words is a hard problem because of sparse context information. Mimicking (Pinter et al., 2017) has been proposed as a solution: given embeddings learned by a standard algorithm, …

FormWord Embeddings