paper-with-me

홈 › Papers

Opening the Black Box of wav2vec Feature Encoder

2022-10-27 · Kwanghee Choi, Eun Jung Yeo

Self-supervised models, namely, wav2vec and its variants, have shown promising results in various downstream tasks in the speech domain. However, their inner workings are poorly understood, calling for in-depth analyses on what the model learns. In this paper, we concentrate on the convolutional feature encoder where its latent space is often speculated to represent discrete acoustic units. To analyze the embedding space in a reductive manner, we feed the synthesized audio signals, which is the summation of simple sine waves. Through extensive experiments, we conclude that various information is embedded inside the feature encoder representations: (1) fundamental frequency, (2) formants, and (3) amplitude, packed with (4) sufficient temporal detail. Further, the information incorporated inside the latent representations is analogous to spectrograms but with a fundamental difference: latent representations construct a metric space so that closer representations imply acoustic similarity.

📄 PDF Abstract BibTeX arXiv:2210.15386

Code (1)

juice500ml/unbox-w2v-convnet 공식 구현 pytorch

Similar Papers 제목 키워드 기반

The Deepfake Detective: Interpreting Neural Forensics Through Sparse Features and Manifolds

2025-12-25 · Subramanyam Sahoo, Jared Junkin arxiv

Deepfake detection models have achieved high accuracy in identifying synthetic media, but their decision processes remain largely opaque. In this paper we present a mechanistic interpretability framework for deepfake det…

DeepFake Detection

Machine Learning Algorithms to Predict Chess960 Result and Develop Opening Themes

2023-10-29 · Shreyan Deo, Nishchal Dwivedi

This work focuses on the analysis of Chess 960, also known as Fischer Random Chess, a variant of traditional chess where the starting positions of the pieces are randomized. The study aims to predict the game outcome usi…

Position

Predicting the performance of hybrid ventilation in buildings using a multivariate attention-based biLSTM Encoder-Decoder neural network

2023-02-08 · Gaurav Chaudhary, Hicham Johra, Laurent Georges, Bjørn Austbø

Hybrid ventilation is an energy-efficient solution to provide fresh air for most climates, given that it has a reliable control system. To operate such systems optimally, a high-fidelity control-oriented modesl is requir…

Decoder

Fairer Chess: A Reversal of Two Opening Moves in Chess Creates Balance Between White and Black

2021-08-05 · Steven J. Brams, Mehmet S. Ismail

Unlike tic-tac-toe or checkers, in which optimal play leads to a draw, it is not known whether optimal play in chess ends in a win for White, a win for Black, or a draw. But after White moves first in chess, if Black has…

Training Neural Nets to Achieve Audio-to-Score Translation: Opening the Black-Box

2018-10-22 · Anonymous

It is suggested that the task of audio-to-score translation offers an adequate testbed to investigate the division of labor between background knowledge and machine learning in the domain of audio pattern recognition, wi…

Translation