paper-with-me

홈 › Papers

Tracking the emergence of linguistic structure in self-supervised models learning from speech

2026-04-02 · Marianne de Heer Kloots, Martijn Bentum, Hosein Mohebbi, Charlotte Pouw, Gaofei Shen, Willem Zuidema arxiv

Self-supervised speech models learn effective representations of spoken language, which have been shown to reflect various aspects of linguistic structure. But when does such structure emerge in model training? We study the encoding of a wide range of linguistic structures, across layers and intermediate checkpoints of six Wav2Vec2 and HuBERT models trained on spoken Dutch. We find that different levels of linguistic structure show notably distinct layerwise patterns as well as learning trajectories, which can partially be explained by differences in their degree of abstraction from the acoustic signal and the timescale at which information from the input is integrated. Moreover, we find that the level at which pre-training objectives are defined strongly affects both the layerwise organization and the learning trajectories of linguistic structures, with greater parallelism induced by higher-order prediction tasks (i.e. iteratively refined pseudo-labels).

📄 PDF Abstract BibTeX arXiv:2604.02043

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unsupervised Chunking with Hierarchical RNN

2023-09-10 · Zijun Wu, Anup Anand Deshmukh, Yongkang Wu, Jimmy Lin 외

In Natural Language Processing (NLP), predicting linguistic structures, such as parsing and chunking, has mostly relied on manual annotations of syntactic structures. This paper introduces an unsupervised approach to chu…

ChunkingSentence

Emergence of Human-Like Attention in Self-Supervised Vision Transformers: an eye-tracking study

2024-10-30 · Takuto Yamamoto, Hirosato Akahoshi, Shigeru Kitazawa

Many models of visual attention have been proposed so far. Traditional bottom-up models, like saliency models, fail to replicate human gaze patterns, and deep gaze prediction models lack biological plausibility due to th…

Gaze Prediction

Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining

2025-09-05 · Deniz Bayazit, Aaron Mueller, Antoine Bosselut arxiv

Large language models (LLMs) learn non-trivial abstractions during pretraining, such as detecting irregular plural noun subjects. However, because traditional evaluation methods (e.g., benchmarking) fail to reveal how mo…

Representation Learning

MAST: A Memory-Augmented Self-supervised Tracker

2020-02-18 · CVPR 2020 6 · Zihang Lai, Erika Lu, Weidi Xie

Recent interest in self-supervised dense tracking has yielded rapid progress, but performance still remains far from supervised methods. We propose a dense tracking model trained on videos without any annotations that su…

Semantic SegmentationSemi-Supervised Video Object SegmentationUnsupervised Video Object SegmentationVideo Object Segmentation+1

Patch2Self: Denoising Diffusion MRI with Self-Supervised Learning​

2020-12-01 · NeurIPS 2020 12 · Shreyas Fadnavis, Joshua Batson, Eleftherios Garyfallidis

Diffusion-weighted magnetic resonance imaging (DWI) is the only non-invasive method for quantifying microstructure and reconstructing white-matter pathways in the living human brain. Fluctuations from multiple sources cr…

DenoisingDiffusion MRISelf-Supervised Learning