paper-with-me

Papers

Example-Based Framework for Perceptually Guided Audio Texture Generation

2023-08-23 · Purnima Kamath, Chitralekha Gupta, Lonce Wyse, Suranga Nanayakkara

Controllable generation using StyleGANs is usually achieved by training the model using labeled data. For audio textures, however, there is currently a lack of large semantically labeled datasets. Therefore, to control generation, we develop a method for semantic control over an unconditionally trained StyleGAN in the absence of such labeled datasets. In this paper, we propose an example-based framework to determine guidance vectors for audio texture generation based on user-defined semantic attributes. Our approach leverages the semantically disentangled latent space of an unconditionally trained StyleGAN. By using a few synthetic examples to indicate the presence or absence of a semantic attribute, we infer the guidance vectors in the latent space of the StyleGAN to control that attribute during generation. Our results show that our framework can find user-defined and perceptually relevant guidance vectors for controllable generation for audio textures. Furthermore, we demonstrate an application of our framework to other tasks, such as selective semantic attribute transfer.

📄 PDF Abstract BibTeX arXiv:2308.11859

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeTexture Synthesis

Methods 이 논문이 사용한 방법론

Adaptive Instance Normalization 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
HuMan(Expedia)||How do I get a human at Expedia? How do I get a human at Expedia? How Do I Get a Human at Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Real-Time Help & Exclusive…
R1 Regularization R_INLINE_MATH_1 Regularization is a regularization technique and gradient penalty for training [generative adversarial…
Feedforward Network A Feedforward Network, or a Multilayer Perceptron (MLP), is a neural network with solely densely connected layers. This is the classic neural network architecture of the…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
StyleGAN 설명 없음

Similar Papers 제목 키워드 기반

“Style” Transfer for Musical Audio Using Multiple Time-Frequency Representations

2018-01-01 · ICLR 2018 1 · Shaun Barry, Youngmoo Kim

Neural Style Transfer has become a popular technique for generating images of distinct artistic styles using convolutional neural networks. This recent success in image style transfer has raised the question of whether s…

Style TransferTexture Synthesis

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

2026-06-03 · Eugene Kwek, Feng Liu, Rui Zhang, Wenpeng Yin arxiv

Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance reconstruction quality with token effic…

Voice Conversion

JSQA: Speech Quality Assessment with Perceptually-Inspired Contrastive Pretraining Based on JND Audio Pairs

2025-07-15 · Junyi Fan, Donald Williamson arxiv

Speech quality assessment (SQA) is often used to learn a mapping from a high-dimensional input space to a scalar that represents the mean opinion score (MOS) of the perceptual speech quality. Learning such a mapping is c…

Contrastive Learning

VGFlow: Visibility guided Flow Network for Human Reposing

2022-11-13 · CVPR 2023 1 · Rishabh Jain, Krishna Kumar Singh, Mayur Hemani, Jingwan Lu 외

The task of human reposing involves generating a realistic image of a person standing in an arbitrary conceivable pose. There are multiple difficulties in generating perceptually accurate images, and existing methods suf…

DiversitySSIM

WaveGuard: Understanding and Mitigating Audio Adversarial Examples

2021-03-04 · Shehzeen Hussain, Paarth Neekhara, Shlomo Dubnov, Julian McAuley 외

There has been a recent surge in adversarial attacks on deep learning based automatic speech recognition (ASR) systems. These attacks pose new challenges to deep learning security and have raised significant concerns in …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition