Localizing Spatial Information in Neural Spatiospectral Filters
Beamforming for multichannel speech enhancement relies on the estimation of spatial characteristics of the acoustic scene. In its simplest form, the delay-and-sum beamformer (DSB) introduces a time delay to all channels to align the desired signal components for constructive superposition. Recent investigations of neural spatiospectral filtering revealed that these filters can be characterized by a beampattern similar to one of traditional beamformers, which shows that artificial neural networks can learn and explicitly represent spatial structure. Using the Complex-valued Spatial Autoencoder (COSPA) as an exemplary neural spatiospectral filter for multichannel speech enhancement, we investigate where and how such networks represent spatial information. We show via clustering that for COSPA the spatial information is represented by the features generated by a gated recurrent unit (GRU) layer that has access to all channels simultaneously and that these features are not source -- but only direction of arrival-dependent.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech EnhancementMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Spatially constrained vs. unconstrained filtering in neural spatiospectral filters for multichannel speech enhancement
When using artificial neural networks for multichannel speech enhancement, filtering is often achieved by estimating a complex-valued mask that is applied to all or one reference channel of the input signal. The estimati…
Speech EnhancementmHealth hyperspectral learning for instantaneous spatiospectral imaging of hemodynamics
Hyperspectral imaging acquires data in both the spatial and frequency domains to offer abundant physical or biological information. However, conventional hyperspectral imaging has intrinsic limitations of bulky instrumen…
compressed sensingAttentional Triple-Encoder Network in Spatiospectral Domains for Medical Image Segmentation
Retinal Optical Coherence Tomography (OCT) segmentation is essential for diagnosing pathology. Traditional methods focus on either spatial or spectral domains, overlooking their combined dependencies. We propose a triple…
Image SegmentationMedical Image SegmentationSemantic SegmentationTowards Localizing Structural Elements: Merging Geometrical Detection with Semantic Verification in RGB-D Data
RGB-D cameras supply rich and dense visual and spatial information for various robotics tasks such as scene understanding, map reconstruction, and localization. Integrating depth and visual information can aid robots in …
3D Plane Detection3d scene graph generationGraph GenerationPanoptic Segmentation+3DeepAf: One-Shot Spatiospectral Auto-Focus Model for Digital Pathology
While Whole Slide Imaging (WSI) scanners remain the gold standard for digitizing pathology samples, their high cost limits accessibility in many healthcare settings. Other low-cost solutions also face critical limitation…
Cancer Classification