paper-with-me

Papers

Scaling Self-Supervised Speech Models Uncovers Deep Linguistic Relationships: Evidence from the Pacific Cluster

2026-03-07 · Minu Kim, Hoirin Kim, David R. Mortensen arxiv

Similarities between language representations derived from Self-Supervised Speech Models (S3Ms) have been observed to primarily reflect geographic proximity or surface typological similarities driven by recent expansion or contact, potentially missing deeper genealogical signals. We investigate how scaling an S3M-based language identification system from 126 to 4,017 languages reshapes this topology, and find a non-linear effect: phylogenetic recovery stays flat up to the 1K scale, but the 4K model undergoes a qualitative shift, resolving both clear lineages and long-term linguistic contact. Most strikingly, a robust Pacific macro-cluster emerges, grouping genealogically unrelated Papuan, Oceanic, and Australian languages, and we trace its driver to a concentrated encoding that captures shared acoustic signatures such as global energy dynamics. These results suggest that massive S3Ms internalize multiple layers of language history, offering a promising perspective for computational phylogenetics and the study of language contact.

📄 PDF Abstract BibTeX arXiv:2603.07238

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identification

Similar Papers 제목 키워드 기반

Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations

2025-09-19 · Linyang He, Qiaolin Wang, Xilin Jiang, Nima Mesgarani arxiv

Transformer-based speech language models (SLMs) have significantly improved neural speech recognition and understanding. While existing research has examined how well SLMs encode shallow acoustic and phonetic features, t…

Self-Supervised LearningSpeech Recognition

SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?

2023-06-14 · Takanori Ashihara, Takafumi Moriya, Kohei Matsuura, Tomohiro Tanaka 외

Self-supervised learning (SSL) for speech representation has been successfully applied in various downstream tasks, such as speech and speaker recognition. More recently, speech SSL models have also been shown to be bene…

Natural Language UnderstandingSelf-Supervised LearningSpeaker RecognitionSpoken Language Understanding

What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training

2025-06-01 · Marianne de Heer Kloots, Hosein Mohebbi, Charlotte Pouw, Gaofei Shen 외

How language-specific are speech representations learned by self-supervised models? Existing work has shown that a range of linguistic features can be successfully decoded from end-to-end models trained only on speech re…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT

2024-09-16 · Ryota Komatsu, Takahiro Shinozaki

Self-supervised speech representation learning has become essential for extracting meaningful features from untranscribed audio. Recent advances highlight the potential of deriving discrete symbols from the features corr…

Acoustic Unit DiscoveryClusteringData AugmentationRepresentation Learning+3

Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations

2025-03-15 · Xue Jiang, Xiulian Peng, Yuan Zhang, Yan Lu

Current large speech language models are mainly based on semantic tokens from discretization of self-supervised learned representations and acoustic tokens from a neural codec, following a semantic-modeling and acoustic-…