paper-with-me

홈 › Papers

Symbolic Autoencoding for Self-Supervised Sequence Learning

2024-02-16 · Mohammad Hossein Amani, Nicolas Mario Baldwin, Amin Mansouri, Martin Josifoski, Maxime Peyrard, Robert West

Traditional language models, adept at next-token prediction in text sequences, often struggle with transduction tasks between distinct symbolic systems, particularly when parallel data is scarce. Addressing this issue, we introduce \textit{symbolic autoencoding} ($\Sigma$AE), a self-supervised framework that harnesses the power of abundant unparallel data alongside limited parallel data. $\Sigma$AE connects two generative models via a discrete bottleneck layer and is optimized end-to-end by minimizing reconstruction loss (simultaneously with supervised loss for the parallel data), such that the sequence generated by the discrete bottleneck can be read out as the transduced input sequence. We also develop gradient-based methods allowing for efficient self-supervised sequence learning despite the discreteness of the bottleneck. Our results demonstrate that $\Sigma$AE significantly enhances performance on transduction tasks, even with minimal parallel data, offering a promising solution for weakly supervised learning scenarios.

📄 PDF Abstract BibTeX arXiv:2402.10575

Code (0)

등록된 구현이 없습니다.

Tasks

Weakly-supervised Learning

Similar Papers 제목 키워드 기반

Word Segmentation on Discovered Phone Units with Dynamic Programming and Self-Supervised Scoring

2022-02-24 · Herman Kamper

Recent work on unsupervised speech segmentation has used self-supervised models with phone and word segmentation modules that are trained jointly. This paper instead revisits an older approach to word segmentation: botto…

Acoustic Unit DiscoverySegmentation

Unsupervised Learning of Neurosymbolic Encoders

2021-07-28 · Eric Zhan, Jennifer J. Sun, Ann Kennedy, Yisong Yue 외

We present a framework for the unsupervised learning of neurosymbolic encoders, which are encoders obtained by composing neural networks with symbolic programs from a domain-specific language. Our framework naturally inc…

DecoderProgram SynthesisSports Analytics

Representation Learning for Sequence Data with Deep Autoencoding Predictive Components

2020-10-07 · ICLR 2021 1 · Junwen Bai, Weiran Wang, Yingbo Zhou, Caiming Xiong

We propose Deep Autoencoding Predictive Components (DAPC) -- a self-supervised representation learning method for sequence data, based on the intuition that useful representations of sequence data should exhibit a simple…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive LearningRepresentation Learning+2

Speech segmentation with a neural encoder model of working memory

2017-09-01 · EMNLP 2017 9 · Micha Elsner, Cory Shain

We present the first unsupervised LSTM speech segmenter as a cognitive model of the acquisition of words from unsegmented input. Cognitive biases toward phonological and syntactic predictability in speech are rooted in t…

Decoder

Robust Human Trajectory Prediction via Self-Supervised Skeleton Representation Learning

2026-02-26 · Taishu Arashima, Hiroshi Kera, Kazuhiko Kawamoto arxiv

Human trajectory prediction plays a crucial role in applications such as autonomous navigation and video surveillance. While recent works have explored the integration of human skeleton sequences to complement trajectory…

Representation LearningTrajectory Prediction