paper-with-me

Papers

Wave to Syntax: Probing spoken language models for syntax

2023-05-30 · Gaofei Shen, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupała

Understanding which information is encoded in deep models of spoken and written language has been the focus of much research in recent years, as it is crucial for debugging and improving these architectures. Most previous work has focused on probing for speaker characteristics, acoustic and phonological information in models of spoken language, and for syntactic information in models of written language. Here we focus on the encoding of syntax in several self-supervised and visually grounded models of spoken language. We employ two complementary probing methods, combined with baselines and reference representations to quantify the degree to which syntactic structure is encoded in the activations of the target models. We show that syntax is captured most prominently in the middle layers of the networks, and more explicitly within models with more parameters.

📄 PDF Abstract BibTeX arXiv:2305.18957

Code (1)

techsword/wave-to-syntax 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

What has LeBenchmark Learnt about French Syntax?

2024-03-04 · Zdravko Dugonjić, Adrien Pupier, Benjamin Lecouteux, Maximin Coavoux

The paper reports on a series of experiments aiming at probing LeBenchmark, a pretrained acoustic model trained on 7k hours of spoken French, for syntactic information. Pretrained acoustic models are increasingly used fo…

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpoken Language Understanding

Audio-Visual Neural Syntax Acquisition

2023-10-11 · Cheng-I Jeff Lai, Freda Shi, Puyuan Peng, Yoon Kim 외

We study phrase structure induction from visually-grounded speech. The core idea is to first segment the speech waveform into sequences of word segments, and subsequently induce phrase structure using the inferred segmen…

Language Acquisition

Vyākarana: A Colorless Green Benchmark for Syntactic Evaluation in Indic Languages

2021-03-01 · EMNLP (MRL) 2021 11 · Rajaswa Patil, Jasleen Dhillon, Siddhant Mahurkar, Saumitra Kulkarni 외

While there has been significant progress towards developing NLU resources for Indic languages, syntactic evaluation has been relatively less explored. Unlike English, Indic languages have rich morphosyntax, grammatical …

Depth EstimationDepth PredictionPOSPOS Tagging+2

The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling

2020-11-23 · Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé, Morgane Rivière 외

We introduce a new unsupervised task, spoken language modeling: the learning of linguistic representations from raw audio signals without any labels, along with the Zero Resource Speech Benchmark 2021: a suite of 4 black…

ClusteringLanguage ModelingLanguage ModellingRepresentation Learning

Deep Clustering of Text Representations for Supervision-free Probing of Syntax

2020-10-24 · Vikram Gupta, Haoyue Shi, Kevin Gimpel, Mrinmaya Sachan

We explore deep clustering of text representations for unsupervised model interpretation and induction of syntax. As these representations are high-dimensional, out-of-the-box methods like KMeans do not work well. Thus, …

ClusteringDeep ClusteringTAG