paper-with-me

Papers

Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation

2024-09-20 · Matthew Caren, Kartik Chandra, Joshua B. Tenenbaum, Jonathan Ragan-Kelley, Karima Ma

We present a method for automatically producing human-like vocal imitations of sounds: the equivalent of "sketching," but for auditory rather than visual representation. Starting with a simulated model of the human vocal tract, we first try generating vocal imitations by tuning the model's control parameters to make the synthesized vocalization match the target sound in terms of perceptually-salient auditory features. Then, to better match human intuitions, we apply a cognitive theory of communication to take into account how human speakers reason strategically about their listeners. Finally, we show through several experiments and user studies that when we add this type of communicative reasoning to our method, it aligns with human intuitions better than matching auditory features alone does. This observation has broad implications for the study of depiction in computer graphics.

📄 PDF Abstract BibTeX arXiv:2409.13507

Code (1)

matthewcaren/vocal-imitation 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Adversarial speech for voice privacy protection from Personalized Speech generation

2024-01-22 · Shihao Chen, Liping Chen, Jie Zhang, KongAik Lee 외

The rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human list…

Speaker Verificationtext-to-speechText to SpeechVoice Conversion

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds

2025-05-30 · Minsu Kang, Seolhee Lee, Choonghyeon Lee, Namhyun Cho

Human to non-human voice conversion (H2NH-VC) transforms human speech into animal or designed vocalizations. Unlike prior studies focused on dog-sounds and 16 or 22.05kHz audio transformation, this work addresses a broad…

Voice Conversion

Vowels and Prosody Contribution in Neural Network Based Voice Conversion Algorithm with Noisy Training Data

2020-03-10

This research presents a neural network based voice conversion (VC) model. While it is a known fact that voiced sounds and prosody are the most important component of the voice conversion framework, what is not known is …

Voice Conversion

A study for the effect of the Emphaticness and language and dialect for Voice Onset Time (VOT) in Modern Standard Arabic (MSA)

2013-05-13 · Sulaiman S. AlDahri

The signal sound contains many different features, including Voice Onset Time (VOT), which is a very important feature of stop sounds in many languages. The only application of VOT values is stopping phoneme subsets. Thi…

SoundStorm: Efficient Parallel Audio Generation

2023-05-16 · Zalán Borsos, Matt Sharifi, Damien Vincent, Eugene Kharitonov 외

We present SoundStorm, a model for efficient, non-autoregressive audio generation. SoundStorm receives as input the semantic tokens of AudioLM, and relies on bidirectional attention and confidence-based parallel decoding…

Audio Generation