paper-with-me

Papers

Multi-Dimensional and Multi-Scale Modeling for Speech Separation Optimized by Discriminative Learning

2023-03-07 · Zhaoxi Mu, Xinyu Yang, Wenjing Zhu

Transformer has shown advanced performance in speech separation, benefiting from its ability to capture global features. However, capturing local features and channel information of audio sequences in speech separation is equally important. In this paper, we present a novel approach named Intra-SE-Conformer and Inter-Transformer (ISCIT) for speech separation. Specifically, we design a new network SE-Conformer that can model audio sequences in multiple dimensions and scales, and apply it to the dual-path speech separation framework. Furthermore, we propose Multi-Block Feature Aggregation to improve the separation effect by selectively utilizing information from the intermediate blocks of the separation network. Meanwhile, we propose a speaker similarity discriminative loss to optimize the speech separation model to address the problem of poor performance when speakers have similar voices. Experimental results on the benchmark datasets WSJ0-2mix and WHAM! show that ISCIT can achieve state-of-the-art results.

📄 PDF Abstract BibTeX arXiv:2303.03737

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis

2024-06-16 · Xuehao Zhou, Mingyang Zhang, Yi Zhou, Zhizheng Wu 외

Generating speech across different accents while preserving speaker identity is crucial for various real-world applications. However, accurately and independently modeling both speaker and accent characteristics in text-…

DisentanglementSpeech Synthesistext-to-speechText to Speech+1

Towards Multi-Scale Style Control for Expressive Speech Synthesis

2021-04-08 · Xiang Li, Changhe Song, Jingbei Li, Zhiyong Wu 외

This paper introduces a multi-scale speech style modeling method for end-to-end expressive speech synthesis. The proposed method employs a multi-scale reference encoder to extract both the global-scale utterance-level an…

Expressive Speech SynthesisSpeech SynthesisStyle Transfer

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

2026-04-23 · Srija Anand, Ashwin Sankar, Ishvinder Sethi, Aaditya Pareek 외 arxiv

Crowdsourced pairwise evaluation has emerged as a scalable approach for assessing foundation models. However, applying it to Text to Speech(TTS) introduces high variance due to linguistic diversity and multidimensional n…

Text to Speech

MEBM-Speech: Multi-scale Enhanced BrainMagic for Robust MEG Speech Detection

2026-02-27 · Li Songyi, Zheng Linze, Liang Jinghua, Zhang Zifeng arxiv

We propose MEBM-Speech, a multi-scale enhanced neural decoder for speech activity detection from non-invasive magnetoencephalography (MEG) signals. Built upon the BrainMagic backbone, MEBM-Speech integrates three complem…

Representation LearningActivity Detection

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling

2025-07-25 · Rongkun Xue, Yazhe Niu, Shuai Hu, Zixin Yin 외 arxiv

Discrete speech tokenization is a fundamental component in speech codecs. However, in large-scale speech-to-speech systems, the complexity of parallel streams from multiple quantizers and the computational cost of high-t…