paper-with-me

Papers

How Generative Spoken Language Modeling Encodes Noisy Speech: Investigation from Phonetics to Syntactics

2023-06-01 · Joonyong Park, Shinnosuke Takamichi, Tomohiko Nakamura, Kentaro Seki, Detai Xin, Hiroshi Saruwatari

We examine the speech modeling potential of generative spoken language modeling (GSLM), which involves using learned symbols derived from data rather than phonemes for speech analysis and synthesis. Since GSLM facilitates textless spoken language processing, exploring its effectiveness is critical for paving the way for novel paradigms in spoken-language processing. This paper presents the findings of GSLM's encoding and decoding effectiveness at the spoken-language and speech levels. Through speech resynthesis experiments, we revealed that resynthesis errors occur at the levels ranging from phonology to syntactics and GSLM frequently resynthesizes natural but content-altered speech.

📄 PDF Abstract BibTeX arXiv:2306.00697

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingResynthesis

Similar Papers 제목 키워드 기반

Augmentation Invariant Discrete Representation for Generative Spoken Language Modeling

2022-09-30 · Itai Gat, Felix Kreuk, Tu Anh Nguyen, Ann Lee 외

Generative Spoken Language Modeling research focuses on optimizing speech Language Models (LMs) using raw audio recordings without accessing any textual supervision. Such speech LMs usually operate over discrete units ob…

Language ModelingLanguage ModellingSpeech-to-Speech Translation

Using ASR-Generated Text for Spoken Language Modeling

2022-05-01 · BigScience (ACL) 2022 5 · Nicolas Hervé, Valentin Pelloin, Benoit Favre, Franck Dary 외

This papers aims at improving spoken language modeling (LM) using very large amount of automatically transcribed speech. We leverage the INA (French National Audiovisual Institute) collection and obtain 19GB of text afte…

Language ModelingLanguage Modelling

Text-Free Prosody-Aware Generative Spoken Language Modeling

2021-09-07 · ACL 2022 5 · Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi 외

Speech pre-training has primarily demonstrated efficacy on classification tasks, while its capability of generating novel speech, similar to how GPT-2 can generate coherent paragraphs, has barely been explored. Generativ…

Language ModelingLanguage Modelling

Confusion2vec 2.0: Enriching Ambiguous Spoken Language Representations with Subwords

2021-02-03 · Prashanth Gurunath Shivakumar, Panayiotis Georgiou, Shrikanth Narayanan

Word vector representations enable machines to encode human language for spoken language understanding and processing. Confusion2vec, motivated from human speech production and perception, is a word vector representation…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Intent DetectionNatural Language Understanding+4

Multimodal neural pronunciation modeling for spoken languages with logographic origin

2018-09-12 · EMNLP 2018 10 · Minh Nguyen, Gia H. Ngo, Nancy F. Chen

Graphemes of most languages encode pronunciation, though some are more explicit than others. Languages like Spanish have a straightforward mapping between its graphemes and phonemes, while this mapping is more convoluted…