paper-with-me

홈 › Papers

Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text

2026-08-24 · Hyeonyu Kim, Hwayeon Kim, Youngwon Choi, Myeongkyun Cho, Huu-Kim Nguyen arxiv

Spoken Language Models (SLMs) generate textual responses directly from speech, offering an alternative to cascaded systems. Despite recent advances, existing SLMs still exhibit weaker instruction-following behavior and limited generalization across diverse tasks compared to text-based language models. Our analysis shows that speech and text representations in current SLMs remain weakly aligned despite strong downstream performance, indicating that structural differences between continuous, temporally varying speech and discrete text remain insufficiently addressed. To address this, we propose a simple framework that decouples length mismatch from semantic alignment and encourages closer correspondence between speech and text representations. Experiments across multiple benchmarks demonstrate competitive performance against strong baselines, underscoring the importance of explicitly addressing structural differences between speech and text in SLM training. Our code is publicly available at https://github.com/jaykim9870/Do_SLMs_Hear_Speech_as_They_Read_Text.

📄 PDF Abstract BibTeX arXiv:2608.22908

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Visual Speech Recognition

2014-09-03 · Ahmad B. A. Hassanat

Lip reading is used to understand or interpret speech without hearing it, a technique especially mastered by people with hearing difficulties. The ability to lip read enables a person with a hearing impairment to communi…

Audio-Visual Speech RecognitionLip Readingobject-detectionObject Detection+5

The Munich Biovoice Corpus: Effects of Physical Exercising, Heart Rate, and Skin Conductance on Human Speech Production

2014-05-01 · LREC 2014 5 · Bj{\"o}rn Schuller, Felix Friedmann, Florian Eyben

We introduce a spoken language resource for the analysis of impact that physical exercising has on human speech production. In particular, the database provides heart rate and skin conductance measurement information alo…

Binary ClassificationHeart rate estimation

Towards Estimating the Upper Bound of Visual-Speech Recognition: The Visual Lip-Reading Feasibility Database

2017-04-26 · Adriana Fernandez-Lopez, Oriol Martinez, Federico M. Sukno

Speech is the most used communication method between humans and it involves the perception of auditory and visual channels. Automatic speech recognition focuses on interpreting the audio signals, although the video can p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Lip Readingspeech-recognition+2

"I've Heard of You!": Generate Spoken Named Entity Recognition Data for Unseen Entities

2024-12-26 · Jiawei Yu, Xiang Geng, Yuang Li, Mengxin Ren 외

Spoken named entity recognition (NER) aims to identify named entities from speech, playing an important role in speech processing. New named entities appear every day, however, annotating their Spoken NER data is costly.…

Domain AdaptationLanguage ModelingLanguage ModellingLarge Language Model+6

A Novel Interpretable and Generalizable Re-synchronization Model for Cued Speech based on a Multi-Cuer Corpus

2023-06-05 · Lufei Gao, Shan Huang, Li Liu

Cued Speech (CS) is a multi-modal visual coding system combining lip reading with several hand cues at the phonetic level to make the spoken language visible to the hearing impaired. Previous studies solved asynchronous …

Lip Reading