paper-with-me

홈 › Papers

Bridging the Gap: Using Deep Acoustic Representations to Learn Grounded Language from Percepts and Raw Speech

2021-12-27 · Gaoussou Youssouf Kebe, Luke E. Richards, Edward Raff, Francis Ferraro, Cynthia Matuszek

Learning to understand grounded language, which connects natural language to percepts, is a critical research area. Prior work in grounded language acquisition has focused primarily on textual inputs. In this work we demonstrate the feasibility of performing grounded language acquisition on paired visual percepts and raw speech inputs. This will allow interactions in which language about novel tasks and environments is learned from end users, reducing dependence on textual inputs and potentially mitigating the effects of demographic bias found in widely available speech recognition systems. We leverage recent work in self-supervised speech representation models and show that learned representations of speech can make language grounding systems more inclusive towards specific groups while maintaining or even increasing general performance.

📄 PDF Abstract BibTeX arXiv:2112.13758

Code (0)

등록된 구현이 없습니다.

Tasks

Language Acquisitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Neural approaches to spoken content embedding

2023-08-28 · Shane Settle

Comparing spoken segments is a central operation to speech processing. Traditional approaches in this area have favored frame-level dynamic programming algorithms, such as dynamic time warping, because they require no su…

Automatic Speech RecognitionDynamic Time Warpingspeech-recognitionSpeech Recognition+1

ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools -- From Consensus Learning to Ambiguity-Driven Emotion Reasoning

2026-02-13 · Esther Sun, Bo-Hao Su, Abinay Reddy Naini, Shinji Watanabe 외 arxiv

Speech Large Language Models (SLLMs) enable high-level emotion reasoning but often produce ungrounded, text-biased judgments without verifiable acoustic evidence. In contrast, self-supervised speech encoders such as WavL…

Emotion Recognition

New Insights into Optimal Alignment of Acoustic and Linguistic Representations for Knowledge Transfer in ASR

2025-09-06 · Xugang Lu, Peng Shen, Hisashi Kawai arxiv

Aligning acoustic and linguistic representations is a central challenge to bridge the pre-trained models in knowledge transfer for automatic speech recognition (ASR). This alignment is inherently structured and asymmetri…

Speech Recognition

Multilingual Jointly Trained Acoustic and Written Word Embeddings

2020-06-24 · Yushi Hu, Shane Settle, Karen Livescu

Acoustic word embeddings (AWEs) are vector representations of spoken word segments. AWEs can be learned jointly with embeddings of character sequences, to generate phonetically meaningful embeddings of written words, or …

Dynamic Time WarpingRetrievalWord Embeddings

Language with Vision: a Study on Grounded Word and Sentence Embeddings

2022-06-17 · Hassan Shahmohammadi, Maria Heitmeier, Elnaz Shafaei-Bajestan, Hendrik P. A. Lensch 외

Grounding language in vision is an active field of research seeking to construct cognitively plausible word and sentence representations by incorporating perceptual knowledge from vision into text-based representations. …

SentenceSentence EmbeddingsVisual GroundingWord Embeddings+1