paper-with-me

홈 › Papers

Structural Bottlenecks on Frequency Representation in End-to-End Audio Models

2026-07-09 · Nicole Cosme-Clifford arxiv

End-to-end neural audio models achieve high-fidelity compression and generation. We might read that performance as evidence they directly represent interpretable features such as pitch and timbre, but a model can produce plausible outputs without doing so. A model may encode these features in any reachable basis, but regardless of which, the features are well described as compositions of time-frequency-localized primitives. Whether state-of-the-art encoders preserve access to these primitives, and thus to compositions of them, remains unclear. Through theoretical analysis and controlled experiments, we show that several state-of-the-art strided convolutional encoders impose two structural bottlenecks, both predictable from architecture and signal structure, on access to these primitives: (1) they collapse primitives into alias equivalence classes, establishing a bound on representational capacity, and (2) they limit the frequency resolution available to learned filters, restricting separability. For well structured data, we find collapse rates of 31-35% and filter bandwidths 10-35x above the theoretical resolution bound, confirming that both bottlenecks arise under realistic signal conditions. We then introduce Gabor Latent Refactorization (GLRF), a lightweight post-hoc intervention that re-expresses encoder latents in a frequency-localized basis, reducing filter bandwidths from 10-35x to 1.5-3x of the theoretical resolution bound while preserving reconstruction fidelity and improving control over attributes like pitch. These results show that the encoders in question predictably degrade access to frequency-localized primitives, entangling the features that depend on them, and that a lightweight, retraining-free intervention can recover much of that access, improving steerability and interpretability.

📄 PDF Abstract BibTeX arXiv:2607.08545

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation

2025-09-23 · Yunzhe Shen, Kai Peng, Leiye Liu, Wei Ji 외 arxiv

Audio-visual segmentation (AVS) plays a critical role in multimodal machine learning by effectively integrating audio and visual cues to precisely segment objects or regions within visual scenes. Recent AVS methods have …

GraFPrint: A GNN-Based Approach for Audio Identification

2024-10-14 · Aditya Bhattacharjee, Shubhr Singh, Emmanouil Benetos

This paper introduces GraFPrint, an audio identification framework that leverages the structural learning capabilities of Graph Neural Networks (GNNs) to create robust audio fingerprints. Our method constructs a k-neares…

Learning De-identified Representations of Prosody from Raw Audio

2021-07-17 · Jack Weston, Raphael Lenain, Udeepa Meepegama, Emil Fristed

We propose a method for learning de-identified prosody representations from raw audio using a contrastive self-supervised signal. Whereas prior work has relied on conditioning models on bottlenecks, we introduce a set of…

Spoken Language Understanding

Structural Causal Bottleneck Models

2026-03-09 · Simon Bing, Jonas Wahl, Jakob Runge arxiv

We introduce structural causal bottleneck models (SCBMs), a novel class of structural causal models. At the core of SCBMs lies the assumption that causal effects between high-dimensional variables only depend on low-dime…

Representation LearningTransfer Learning

SUPER Decoder Block for Reconstruction-Aware U-Net Variants

2025-11-14 · Siheon Joo, Hongjo Kim arxiv

Skip-connected encoder-decoder architectures (U-Net variants) are widely adopted for inverse problems but still suffer from information loss, limiting recovery of fine high-frequency details. We present Selectively Suppr…

Crack SegmentationImage Denoising