paper-with-me

Papers

GmSLM : Generative Marmoset Spoken Language Modeling

2025-09-11 · Talia Sternberg, Michael London, David Omer, Yossi Adi arxiv

Marmoset monkeys exhibit complex vocal communication, challenging the view that nonhuman primates vocal communication is entirely innate, and show similar features of human speech, such as vocal labeling of others and turn-taking. Studying their vocal communication offers a unique opportunity to link it with brain activity-especially given the difficulty of accessing the human brain in speech and language research. Since Marmosets communicate primarily through vocalizations, applying standard LLM approaches is not straightforward. We introduce Generative Marmoset Spoken Language Modeling (GmSLM), an optimized spoken language model pipeline for Marmoset vocal communication. We designed a novel zero-shot evaluation metrics using unsupervised in-the-wild data, alongside weakly labeled conversational data, to assess GmSLM and demonstrate its advantage over a basic human-speech-based baseline. GmSLM generated vocalizations closely matched real resynthesized samples acoustically and performed well on downstream tasks. Despite being fully unsupervised, GmSLM effectively distinguish real from artificial conversations and may support further investigations of the neural basis of vocal communication and provides a practical framework linking vocalization and brain activity. We believe GmSLM stands to benefit future work in neuroscience, bioacoustics, and evolutionary biology. Samples are provided under: pages.cs.huji.ac.il/adiyoss-lab/GmSLM.

📄 PDF Abstract BibTeX arXiv:2509.09198

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Transformer Model for Segmentation, Classification, and Caller Identification of Marmoset Vocalization

2024-10-30 · Bin Wu, Shinnosuke Takamichi, Sakriani Sakti, Satoshi Nakamura

Marmoset, a highly vocalized primate, has become a popular animal model for studying social-communicative behavior and its underlying mechanism comparing with human infant linguistic developments. In the study of vocal c…

On feature representations for marmoset vocal communication analysis

2025-04-21 · Eklavya Sarkar, Kaja Wierucka, Alexandra B. Bosshard, Judith Burkart 외

The acoustic analysis of marmoset (Callithrix jacchus) vocalizations is often used to understand the evolutionary origins of human language. Currently, the analysis is largely carried out in a manual or semi-manual manne…

Self-Supervised Learning

How Generative Spoken Language Modeling Encodes Noisy Speech: Investigation from Phonetics to Syntactics

2023-06-01 · Joonyong Park, Shinnosuke Takamichi, Tomohiko Nakamura, Kentaro Seki 외

We examine the speech modeling potential of generative spoken language modeling (GSLM), which involves using learned symbols derived from data rather than phonemes for speech analysis and synthesis. Since GSLM facilitate…

Language ModelingLanguage ModellingResynthesis

Augmentation Invariant Discrete Representation for Generative Spoken Language Modeling

2022-09-30 · Itai Gat, Felix Kreuk, Tu Anh Nguyen, Ann Lee 외

Generative Spoken Language Modeling research focuses on optimizing speech Language Models (LMs) using raw audio recordings without accessing any textual supervision. Such speech LMs usually operate over discrete units ob…

Language ModelingLanguage ModellingSpeech-to-Speech Translation

Text-Free Prosody-Aware Generative Spoken Language Modeling

2021-09-07 · ACL 2022 5 · Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi 외

Speech pre-training has primarily demonstrated efficacy on classification tasks, while its capability of generating novel speech, similar to how GPT-2 can generate coherent paragraphs, has barely been explored. Generativ…

Language ModelingLanguage Modelling