paper-with-me

Papers

Modeling speech recognition and synthesis simultaneously: Encoding and decoding lexical and sublexical semantic information into speech with no access to speech data

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Human speakers encode information into raw speech which is then decoded by the listeners. This complex relationship between encoding (production) and decoding (perception) is often modeled separately. Here, we test how decoding of lexical and sublexical semantic information can emerge automatically from raw speech in unsupervised generative deep convolutional networks that combine both the production and perception principle. We introduce, to our knowledge, the most challenging objective in unsupervised lexical learning: an unsupervised network that must learn to assign unique representations for lexical items with no direct access to training data. We train several models (ciwGAN and fiwGAN by Beguš 2021) and test how the networks classify raw acoustic lexical items in the unobserved test data. Strong evidence in favor of lexical learning emerges. The architecture that combines the production and perception principles is thus able to learn to decode unique information from raw acoustic data in an unsupervised manner without ever accessing real training data. We propose a technique to explore lexical and sublexical learned representations in the classifier network. The results bear implications for both unsupervised speech synthesis and recognition as well as for unsupervised semantic modeling as language models increasingly bypass text and operate from raw acoustics.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionSpeech Synthesis

Similar Papers 제목 키워드 기반

Modeling speech recognition and synthesis simultaneously: Encoding and decoding lexical and sublexical semantic information into speech with no direct access to speech data

2022-03-22 · Gašper Beguš, Alan Zhou

Human speakers encode information into raw speech which is then decoded by the listeners. This complex relationship between encoding (production) and decoding (perception) is often modeled separately. Here, we test how e…

speech-recognitionSpeech RecognitionSpeech Synthesis

DeepTalk: Vocal Style Encoding for Speaker Recognition and Speech Synthesis

2020-12-09 · Anurag Chowdhury, Arun Ross, Prabu David

Automatic speaker recognition algorithms typically characterize speech audio using short-term spectral features that encode the physiological and anatomical aspects of speech production. Such algorithms do not fully capi…

Speaker RecognitionSpeech Synthesis

How Generative Spoken Language Modeling Encodes Noisy Speech: Investigation from Phonetics to Syntactics

2023-06-01 · Joonyong Park, Shinnosuke Takamichi, Tomohiko Nakamura, Kentaro Seki 외

We examine the speech modeling potential of generative spoken language modeling (GSLM), which involves using learned symbols derived from data rather than phonemes for speech analysis and synthesis. Since GSLM facilitate…

Language ModelingLanguage ModellingResynthesis

Towards Joint Modeling of Dialogue Response and Speech Synthesis based on Large Language Model

2023-09-20 · Xinyu Zhou, Delong Chen, Yudong Chen

This paper explores the potential of constructing an AI spoken dialogue system that "thinks how to respond" and "thinks how to speak" simultaneously, which more closely aligns with the human speech production process com…

ChatbotLanguage ModelingLanguage ModellingLarge Language Model+4

Pronunciation-aware unique character encoding for RNN Transducer-based Mandarin speech recognition

2022-07-29 · Peng Shen, Xugang Lu, Hisashi Kawai

For Mandarin end-to-end (E2E) automatic speech recognition (ASR) tasks, compared to character-based modeling units, pronunciation-based modeling units could improve the sharing of modeling units in model training but mee…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition