paper-with-me

Papers Target Sound Extraction

“Target Sound Extraction” 태그가 달린 논문 18편 · 필터 해제

SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios

2026-07-02 · Ziyang Jiang, Yu Chen, Zexu Pan, Xinyuan Qian 외 arxiv

Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remains challenging for current deep learning-based systems. Sound source localization …

Sound Source LocalizationTarget Sound Extraction

SoundCompass: Navigating Target Sound Extraction With Effective Directional Clue Integration In Complex Acoustic Scenes

2025-09-23 · Dayun Choi, Jung-Woo Choi arxiv

Recent advances in target sound extraction (TSE) utilize directional clues derived from direction of arrival (DoA), which represent an inherent spatial property of sound available in any acoustic scene. However, previous…

Target Sound Extraction

SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction

2025-05-30 · Tuochao Chen, D Shin, Hakan Erdogan, Sinan Hersek

This paper introduces SoundSculpt, a neural network designed to extract target sound fields from ambisonic recordings. SoundSculpt employs an ambisonic-in-ambisonic-out architecture and is conditioned on both spatial inf…

Image SegmentationSemantic SegmentationTarget Sound Extraction

Leveraging Audio-Only Data for Text-Queried Target Sound Extraction

2024-09-20 · Kohei Saijo, Janek Ebbers, François G. Germain, Sameer Khurana 외

The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access to large-scale text-audio pairs to addre…

Target Sound Extraction

Multichannel-to-Multichannel Target Sound Extraction Using Direction and Timestamp Clues

2024-09-19 · Dayun Choi, Jung-Woo Choi

We propose a multichannel-to-multichannel target sound extraction (M2M-TSE) framework for separating multichannel target signals from a multichannel mixture of sound sources. Target sound extraction (TSE) isolates a spec…

Inductive BiasTarget Sound Extraction

Language-Queried Target Sound Extraction Without Parallel Training Data

2024-09-14 · Hao Ma, Zhiyuan Peng, Xu Li, Yukai Li 외

Language-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extensively annotated parallel audio-text data…

Language ModellingLarge Language ModelRetrievalTarget Sound Extraction

SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer

2024-09-12 · Helin Wang, Jiarui Hai, Yen-Ju Lu, Karan Thakkar 외

In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-c…

Target Sound Extraction

Cross-attention Inspired Selective State Space Models for Target Sound Extraction

2024-09-07 · Donghang Wu, Yiwen Wang, Xihong Wu, Tianshu Qu

The Transformer model, particularly its cross-attention module, is widely used for feature fusion in target sound extraction which extracts the signal of interest based on given clues. Despite its effectiveness, this app…

Computational EfficiencyMambaState Space ModelsTarget Sound Extraction

Can all variations within the unified mask-based beamformer framework achieve identical peak extraction performance?

2024-07-22 · Atsuo Hiroe, Katsutoshi Itoyama, Kazuhiro Nakadai

This study investigates mask-based beamformers (BFs), which estimate filters for target sound extraction (TSE) using time-frequency masks. Although multiple mask-based BFs have been proposed, no consensus has been reache…

AllTarget Sound Extraction

CATSE: A Context-Aware Framework for Causal Target Sound Extraction

2024-03-21 · Shrishail Baligar, Mikolaj Kegler, Bryce Irvin, Marko Stamenovic 외

Target Sound Extraction (TSE) focuses on the problem of separating sources of interest, indicated by a user's cue, from the input mixture. Most existing solutions operate in an offline fashion and are not suited to the l…

Target Sound Extraction

CLAPSep: Leveraging Contrastive Pre-trained Model for Multi-Modal Query-Conditioned Target Sound Extraction

2024-02-27 · Hao Ma, Zhiyuan Peng, Xu Li, Mingjie Shao 외

Universal sound separation (USS) aims to extract arbitrary types of sounds from real-world recordings. This can be achieved by language-queried target sound extraction (TSE), which typically consists of two components: a…

Target Sound Extraction

Online Similarity-and-Independence-Aware Beamformer for Low-latency Target Sound Extraction

2023-12-27 · Atsuo Hiroe

This study introduces an online target sound extraction (TSE) process using the similarity-and-independence-aware beamformer (SIBF) derived from an iterative batch algorithm. The study aimed to reduce latency while maint…

blind source separationTarget Sound Extraction

Semantic Hearing: Programming Acoustic Scenes with Binaural Hearables

2023-11-01 · Bandhav Veluri, Malek Itani, Justin Chan, Takuya Yoshioka 외

Imagine being able to listen to the birds chirping in a park without hearing the chatter from other hikers, or being able to block out traffic noise on a busy street while still being able to hear emergency sirens and ca…

Target Sound Extraction

DPM-TSE: A Diffusion Probabilistic Model for Target Sound Extraction

2023-10-06 · Jiarui Hai, Helin Wang, Dongchao Yang, Karan Thakkar 외

Common target sound extraction (TSE) approaches primarily relied on discriminative approaches in order to separate the target sound while minimizing interference from the unwanted sources, with varying success in separat…

Target Sound Extraction

Target Sound Extraction with Variable Cross-modality Clues

2023-03-15 · Chenda Li, Yao Qian, Zhuo Chen, Dongmei Wang 외

Automatic target sound extraction (TSE) is a machine learning approach to mimic the human auditory perception capability of attending to a sound source of interest from a mixture of sources. It often uses a model conditi…

AudioCapsTarget Sound Extraction

Real-Time Target Sound Extraction

2022-11-04 · Bandhav Veluri, Justin Chan, Malek Itani, Tuochao Chen 외

We present the first neural network model to achieve real-time and streaming target sound extraction. To accomplish this, we propose Waveformer, an encoder-decoder architecture with a stack of dilated causal convolution …

DecoderStreaming Target Sound ExtractionTarget Sound Extraction

SoundBeam: Target sound extraction conditioned on sound-class labels and enrollment clues for increased performance and continuous learning

2022-04-08 · Marc Delcroix, Jorge Bennasar Vázquez, Tsubasa Ochiai, Keisuke Kinoshita 외

In many situations, we would like to hear desired sound events (SEs) while being able to ignore interference. Target sound extraction (TSE) tackles this problem by estimating the audio signal of the sounds of target SE c…

Target Sound Extraction

Few-shot learning of new sound classes for target sound extraction

2021-06-14 · Marc Delcroix, Jorge Bennasar Vázquez, Tsubasa Ochiai, Keisuke Kinoshita 외

Target sound extraction consists of extracting the sound of a target acoustic event (AE) class from a mixture of AE sounds. It can be realized using a neural network that extracts the target sound conditioned on a 1-hot …

Few-Shot LearningTarget Sound Extraction
1–18 / 18