paper-with-me

Papers

Digital Einstein Experience: Fast Text-to-Speech for Conversational AI

2021-07-21 · Joanna Rownicka, Kilian Sprenkamp, Antonio Tripiana, Volodymyr Gromoglasov, Timo P Kunz

We describe our approach to create and deliver a custom voice for a conversational AI use-case. More specifically, we provide a voice for a Digital Einstein character, to enable human-computer interaction within the digital conversation experience. To create the voice which fits the context well, we first design a voice character and we produce the recordings which correspond to the desired speech attributes. We then model the voice. Our solution utilizes Fastspeech 2 for log-scaled mel-spectrogram prediction from phonemes and Parallel WaveGAN to generate the waveforms. The system supports a character input and gives a speech waveform at the output. We use a custom dictionary for selected words to ensure their proper pronunciation. Our proposed cloud architecture enables for fast voice delivery, making it possible to talk to the digital version of Albert Einstein in real-time.

📄 PDF Abstract BibTeX arXiv:2107.10658

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Position-Wise Feed-Forward Layer 설명 없음
FastSpeech 2 FastSpeech2 is a text-to-speech model that aims to improve upon FastSpeech by better solving the one-to-many mapping problem in TTS, i.e., multiple speech variations…
Adam 설명 없음
Phase Shuffle Phase Shuffle is a technique for removing pitched noise artifacts that come from using transposed convolutions in audio generation models. Phase shuffle is an operation with…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

A Platform for Interactive AI Character Experiences

2026-01-03 · Rafael Wampfler, Chen Yang, Dillon Elste, Nikola Kovacevic 외 arxiv

From movie characters to modern science fiction - bringing characters into interactive, story-driven conversations has captured imaginations across generations. Achieving this vision is highly challenging and requires mu…

Prompt Engineering

Einstein World Models

2026-06-25 · Munachiso Samuel Nwadike, Zangir Iklassov, Ali Mekky, Zayd M. Kawakibi Zuhri 외 arxiv

Does intelligence require the ability to reason about phenomena beyond direct experience? It is natural to suspect that some complex thought cannot be captured through language alone. However, of particular concern to th…

ViDA-MAN: Visual Dialog with Digital Humans

2021-10-26 · Tong Shen, Jiawei Zuo, Fan Shi, Jin Zhang 외

We demonstrate ViDA-MAN, a digital-human agent for multi-modal interaction, which offers realtime audio-visual responses to instant speech inquiries. Compared to traditional text or voice-based system, ViDA-MAN offers hu…

speech-recognitionSpeech Recognitiontext-to-speechText to Speech+2

A Planck Radiation and Quantization Scheme for Human Cognition and Language

2022-01-10 · Diederik Aerts, Lester Beltran

As a result of the identification of 'identity' and 'indistinguishability' and strong experimental evidence for the presence of the associated Bose-Einstein statistics in human cognition and language, we argued in previo…

Quantization

Adaptive transmission for radar arrays using Weiss-Weinstein bounds

2018-08-14

We present an algorithm for adaptive selection of pulse repetition frequency or antenna activations for Doppler and DoA estimation. The adaptation is performed sequentially using a Bayesian filter, responsible for updati…