paper-with-me

Papers

SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction

2025-05-30 · Tuochao Chen, D Shin, Hakan Erdogan, Sinan Hersek

This paper introduces SoundSculpt, a neural network designed to extract target sound fields from ambisonic recordings. SoundSculpt employs an ambisonic-in-ambisonic-out architecture and is conditioned on both spatial information (e.g., target direction obtained by pointing at an immersive video) and semantic embeddings (e.g., derived from image segmentation and captioning). Trained and evaluated on synthetic and real ambisonic mixtures, SoundSculpt demonstrates superior performance compared to various signal processing baselines. Our results further reveal that while spatial conditioning alone can be effective, the combination of spatial and semantic information is beneficial in scenarios where there are secondary sound sources spatially close to the target. Additionally, we compare two different semantic embeddings derived from a text description of the target sound using text encoders.

📄 PDF Abstract BibTeX arXiv:2506.00273

Code (0)

등록된 구현이 없습니다.

Tasks

Image SegmentationSemantic SegmentationTarget Sound Extraction

Similar Papers 제목 키워드 기반

Weakly Guided and Autoregressive Beamformer Parameterization for Generalizable Moving Speaker Extraction in Higher-Order Ambisonics

2026-07-05 · Jakob Kienegger, Tal Peer, Sina Khanagha, Timo Gerkmann arxiv

Linear spatial filters (beamformers) enable robust, generalizable and interpretable speech enhancement with performance guarantees under ideal parameterization. Modern beamformers are often parameterized by deep neural n…

Speech Enhancement

Directional emphasis in ambisonics

2018-05-24

We describe an ambisonics enhancement method that increases the signal strength in specified directions at low computational cost. The method can be used in a static setup to emphasize the signal arriving from a particul…

Direction of Arrival Estimation of Noisy Speech Using Convolutional Recurrent Neural Networks with Higher-Order Ambisonics Signals

2021-02-19 · Nils Poschadel, Robert Hupke, Stephan Preihs, Jürgen Peissig

Training convolutional recurrent neural networks on first-order Ambisonics signals is a well-known approach when estimating the direction of arrival for speech/sound signals. In this work, we investigate whether increasi…

Direction of Arrival Estimation

First Order Ambisonics Domain Spatial Augmentation for DNN-based Direction of Arrival Estimation

2019-10-10 · Luca Mazzon, Yuma Koizumi, Masahiro Yasuda, Noboru Harada

In this paper, we propose a novel data augmentation method for training neural networks for Direction of Arrival (DOA) estimation. This method focuses on expanding the representation of the DOA subspace of a dataset. Giv…

Data AugmentationDirection of Arrival Estimation

A first-order DirAC-based parametric Ambisonic coder for immersive communications

2025-03-30 · Guillaume Fuchs, Florin Ghido, Dominik Weckbecker, Oliver Thiergart

Directional Audio Coding (DirAC) is a proven method for parametrically representing a 3D audio scene in B-format and is capable of reproducing it on arbitrary loudspeaker layouts. Although such a method seems well suited…