paper-with-me

Papers

SSR: Alignment-Aware Modality Connector for Speech Language Models

2024-09-30 · Weiting Tan, Hirofumi Inaguma, Ning Dong, Paden Tomasello, Xutai Ma

Fusing speech into pre-trained language model (SpeechLM) usually suffers from inefficient encoding of long-form speech and catastrophic forgetting of pre-trained text modality. We propose SSR-Connector (Segmented Speech Representation Connector) for better modality fusion. Leveraging speech-text alignments, our approach segments and compresses speech features to match the granularity of text embeddings. Additionally, we introduce a two-stage training pipeline that includes the distillation and fine-tuning phases to mitigate catastrophic forgetting. SSR-Connector outperforms existing mechanism for speech-text modality fusion, consistently achieving better speech understanding (e.g., +10 accuracy on StoryCloze and +20 on Speech-MMLU) while preserving pre-trained text ability.

📄 PDF Abstract BibTeX arXiv:2410.00168

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMMLU

Similar Papers 제목 키워드 기반

ChartMoE: Mixture of Expert Connector for Advanced Chart Understanding

2024-09-05 · Zhengzhuo Xu, Bowen Qu, Yiyan Qi, Sinan Du 외

Automatic chart understanding is crucial for content comprehension and document parsing. Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in chart understanding through domain-specific a…

Chart Understanding

Aligning Pre-trained Models for Spoken Language Translation

2024-11-27 · Šimon Sedláček, Santosh Kesiraju, Alexander Polok, Jan Černocký

This paper investigates a novel approach to end-to-end speech translation (ST) based on aligning frozen pre-trained automatic speech recognition (ASR) and machine translation (MT) models via a small connector module (Q-F…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+2

Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries

2026-01-26 · Yuchen Zhang, Ravi Shekhar, Haralambos Mouratidis arxiv

Large Language Model (LLM)-powered Automatic Speech Recognition (ASR) systems achieve strong performance with limited resources by linking a frozen speech encoder to a pretrained LLM via a lightweight connector. Prior wo…

Speech Recognition

Scaling-Aware Adapter for Structure-Grounded LLM Reasoning

2026-02-02 · Zihao Jing, Qiuhao Zeng, Ruiyi Fang, Yan Yi Li 외 arxiv

Large language models (LLMs) are enabling reasoning over 2D and 3D structures, yet existing methods remain modality-specific and typically compress structural inputs through sequence-based tokenization or fixed-length qu…

Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models

2026-01-24 · Toshiki Nakai, Varsha Suresh, Vera Demberg arxiv

Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether this reflects shared computation for the same language or modality-spec…

Text Generation