paper-with-me

Papers

Integrating Form and Meaning: A Multi-Task Learning Model for Acoustic Word Embeddings

2022-09-14 · Badr M. Abdullah, Bernd Möbius, Dietrich Klakow

Models of acoustic word embeddings (AWEs) learn to map variable-length spoken word segments onto fixed-dimensionality vector representations such that different acoustic exemplars of the same word are projected nearby in the embedding space. In addition to their speech technology applications, AWE models have been shown to predict human performance on a variety of auditory lexical processing tasks. Current AWE models are based on neural networks and trained in a bottom-up approach that integrates acoustic cues to build up a word representation given an acoustic or symbolic supervision signal. Therefore, these models do not leverage or capture high-level lexical knowledge during the learning process. In this paper, we propose a multi-task learning model that incorporates top-down lexical knowledge into the training procedure of AWEs. Our model learns a mapping between the acoustic input and a lexical representation that encodes high-level information such as word semantics in addition to bottom-up form-based supervision. We experiment with three languages and demonstrate that incorporating lexical knowledge improves the embedding space discriminability and encourages the model to better separate lexical categories.

📄 PDF Abstract BibTeX arXiv:2209.06633

Code (1)

uds-lsv/semantically_enriched_awes 공식 구현 pytorch

Tasks

FormMulti-Task LearningWord Embeddings

Similar Papers 제목 키워드 기반

Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion

2025-10-01 · Ahmed Adel Attia, Jing Liu, Carol Espy Wilson arxiv

Prior works have investigated the use of articulatory features as complementary representations for automatic speech recognition (ASR), but their use was largely confined to shallow acoustic models. In this work, we revi…

Speech Recognition

Multilingual Extraction and Recognition of Implicit Discourse Relations in Speech and Text

2026-02-04 · Ahmed Ruby, Christian Hardmeier, Sara Stymne arxiv

Implicit discourse relation classification is a challenging task, as it requires inferring meaning from context. While contextual cues can be distributed across modalities and vary across languages, they are not always c…

Relation ClassificationCross-Lingual Transfer

Decoding Insect Song: A Multitask Semisupervised Orthoptera Bioacoustic Classifier

2026-06-11 · Olga Isupova, Danil Kuzin, Ella Browning, Tom Mills 외 arxiv

Passive acoustic monitoring holds great promise for ecological inference, yet existing automated tools are typically narrowly trained and non-transferable. We address these limitations with PULSE, a semi-supervised, mult…

Self-Supervised LearningKnowledge DistillationActive Learning

Importance-based Multimodal Autoencoder

2021-01-01 · Sayan Ghosh, Eugene Laksana, Louis-Philippe Morency, Stefan Scherer

Integrating information from multiple modalities (e.g., verbal, acoustic and visual data) into meaningful representations has seen great progress in recent years. However, two challenges are not sufficiently addressed …

Learning to Detect Novel and Fine-Grained Acoustic Sequences Using Pretrained Audio Representations

2023-05-03 · Vasudha Kowtha, Miquel Espi Marques, Jonathan Huang, Yichi Zhang 외

This work investigates pretrained audio representations for few shot Sound Event Detection. We specifically address the task of few shot detection of novel acoustic sequences, or sound events with semantically meaningful…

Event DetectionFew-Shot LearningSound Event Detection