paper-with-me

Papers

Audio-Visual Neural Syntax Acquisition

2023-10-11 · Cheng-I Jeff Lai, Freda Shi, Puyuan Peng, Yoon Kim, Kevin Gimpel, Shiyu Chang, Yung-Sung Chuang, Saurabhchand Bhati, David Cox, David Harwath, Yang Zhang, Karen Livescu, James Glass

We study phrase structure induction from visually-grounded speech. The core idea is to first segment the speech waveform into sequences of word segments, and subsequently induce phrase structure using the inferred segment-level continuous representations. We present the Audio-Visual Neural Syntax Learner (AV-NSL) that learns phrase structure by listening to audio and looking at images, without ever being exposed to text. By training on paired images and spoken captions, AV-NSL exhibits the capability to infer meaningful phrase structures that are comparable to those derived by naturally-supervised text parsers, for both English and German. Our findings extend prior work in unsupervised language acquisition from speech and grounded grammar induction, and present one approach to bridge the gap between the two topics.

📄 PDF Abstract BibTeX arXiv:2310.07654

Code (0)

등록된 구현이 없습니다.

Tasks

Language Acquisition

Similar Papers 제목 키워드 기반

What is Learned in Visually Grounded Neural Syntax Acquisition

2020-05-04 · ACL 2020 6 · Noriyuki Kojima, Hadar Averbuch-Elor, Alexander M. Rush, Yoav Artzi

Visual features are a promising signal for learning bootstrap textual models. However, blackbox learning models make it difficult to isolate the specific contribution of visual components. In this analysis, we consider t…

The syntax-semantics interface in a child's path: A study of 3- to 11-year-olds' elicited production of Mandarin recursive relative clauses

2024-06-06 · Caimei Yang, Qihang Yang, Xingzhi Su, Chenxi Fu 외

There have been apparently conflicting claims over the syntax-semantics relationship in child acquisition. However, few of them have assessed the child's path toward the acquisition of recursive relative clauses (RRCs). …

Language AcquisitionObject

Mod\'elisation des processus d'acquisition syntaxique par jeux de langage entre agents artificiels (Modeling Syntactic Acquisition by Language Games between Artificial Agents )

2018-05-01 · JEPTALNRECITAL 2018 5 · Marie Marcia, Isabelle Tellier

Dans cet article, nous pr{\'e}sentons une mod{\'e}lisation de la situation d{'}acquisition de la syntaxe de sa langue maternelle par un enfant inspir{\'e}e des {``}jeux de langages{''} de Luc Steels. Le mod{\`e}le suppos…

MAVD: The First Open Large-Scale Mandarin Audio-Visual Dataset with Depth Information

2023-06-04 · Jianrong Wang, Yuchen Huo, Li Liu, Tianyi Xu 외

Audio-visual speech recognition (AVSR) gains increasing attention from researchers as an important part of human-computer interaction. However, the existing available Mandarin audio-visual datasets are limited and lack t…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

A Synchronized Audio-Visual Multi-View Capture System

2026-03-24 · Xiangwei Shi, Gara Dorta, Ruud de Jong, Ojas Shirekar 외 arxiv

Multi-view capture systems have been an important tool in research for recording human motion under controlling conditions. Most existing systems are specified around video streams and provide little or no support for au…

Video Alignment