paper-with-me

Papers

Single-Microphone Speaker Separation and Voice Activity Detection in Noisy and Reverberant Environments

2024-01-07 · Renana Opochinsky, Mordehay Moradi, Sharon Gannot

Speech separation involves extracting an individual speaker's voice from a multi-speaker audio signal. The increasing complexity of real-world environments, where multiple speakers might converse simultaneously, underscores the importance of effective speech separation techniques. This work presents a single-microphone speaker separation network with TF attention aiming at noisy and reverberant environments. We dub this new architecture as Separation TF Attention Network (Sep-TFAnet). In addition, we present a variant of the separation network, dubbed $ \text{Sep-TFAnet}^{\text{VAD}}$, which incorporates a voice activity detector (VAD) into the separation network. The separation module is based on a temporal convolutional network (TCN) backbone inspired by the Conv-Tasnet architecture with multiple modifications. Rather than a learned encoder and decoder, we use short-time Fourier transform (STFT) and inverse short-time Fourier transform (iSTFT) for the analysis and synthesis, respectively. Our system is specially developed for human-robotic interactions and should support online mode. The separation capabilities of $ \text{Sep-TFAnet}^{\text{VAD}}$ and Sep-TFAnet were evaluated and extensively analyzed under several acoustic conditions, demonstrating their advantages over competing methods. Since separation networks trained on simulated data tend to perform poorly on real recordings, we also demonstrate the ability of the proposed scheme to better generalize to realistic examples recorded in our acoustic lab by a humanoid robot. Project page: https://Sep-TFAnet.github.io

📄 PDF Abstract BibTeX arXiv:2401.03448

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity DetectionDecoderSpeaker SeparationSpeech Separation

Similar Papers 제목 키워드 기반

MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild

2026-06-15 · Haotian Qi, Gabriel Skantze arxiv

Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human-robot interaction scenarios. We introduce MuVAP, a causal multimodal framework t…

Improving speaker discrimination of target speech extraction with time-domain SpeakerBeam

2020-01-23 · Marc Delcroix, Tsubasa Ochiai, Katerina Zmolikova, Keisuke Kinoshita 외

Target speech extraction, which extracts a single target source in a mixture given clues about the target speaker, has attracted increasing attention. We have recently proposed SpeakerBeam, which exploits an adaptation u…

Speaker IdentificationSpeech ExtractionSpeech Separation

Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario

2020-05-14 · Ivan Medennikov, Maxim Korenevsky, Tatiana Prisyach, Yuri Khokhlov 외

Speaker diarization for real-life scenarios is an extremely challenging problem. Widely used clustering-based diarization approaches perform rather poorly in such conditions, mainly due to the limited ability to handle o…

Action DetectionActivity DetectionBinary ClassificationClustering+2

Latent Gaussian Activity Propagation: Using Smoothness and Structure to Separate and Localize Sounds in Large Noisy Environments

2018-12-01 · NeurIPS 2018 12 · Daniel Johnson, Daniel Gorelik, Ross E. Mawhorter, Kyle Suver 외

We present an approach for simultaneously separating and localizing multiple sound sources using recorded microphone data. Inspired by topic models, our approach is based on a probabilistic model of inter-microphone phas…

Bayesian InferencePositionTopic Models

Spatial-Temporal Activity-Informed Diarization and Separation

2024-01-30 · Yicheng Hsu, Ssuhan Chen, Mingsian R. Bai

A robust multichannel speaker diarization and separation system is proposed by exploiting the spatio-temporal activity of the speakers. The system is realized in a hybrid architecture that combines the array signal proce…

speaker-diarizationSpeaker DiarizationSpeaker Separation