paper-with-me

홈 › Papers

VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion

2025-09-19 · Dimitrios Damianos, Leon Voukoutis, Georgios Paraskevopoulos, Vassilis Katsouros arxiv

We present a multimodal fusion framework that bridges pre-trained decoder-based large language models (LLM) and acoustic encoder-decoder architectures such as Whisper, with the aim of building speech-enabled LLMs. Instead of directly using audio embeddings, we explore an intermediate audio-conditioned text space as a more effective mechanism for alignment. Our method operates fully in continuous text representation spaces, fusing Whisper's hidden decoder states with those of an LLM through cross-modal attention, and supports both offline and streaming modes. We introduce \textit{VoxKrikri}, the first Greek speech LLM, and show through analysis that our approach effectively aligns representations across modalities. These results highlight continuous space fusion as a promising path for multilingual and low-resource speech LLMs, while achieving state-of-the-art results for Automatic Speech Recognition in Greek, providing an average $\sim20\%$ relative improvement across benchmarks.

📄 PDF Abstract BibTeX arXiv:2509.15667

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Krikri: Advancing Open Large Language Models for Greek

2025-05-19 · Dimitris Roussis, Leon Voukoutis, Georgios Paraskevopoulos, Sokratis Sofianopoulos 외

We introduce Llama-Krikri-8B, a cutting-edge Large Language Model tailored for the Greek language, built on Meta's Llama 3.1-8B. Llama-Krikri-8B has been extensively trained on high-quality Greek data to ensure superior …

Code GenerationLanguage ModelingLanguage ModellingLarge Language Model+1

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

2026-06-10 · Sangmin Lee, Eekgyun Ahn, Woongjib Choi, Hong-Goo Kang arxiv

We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phoneme (G2P)-based approaches are limited to around 100 languages due to …

LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement

2025-03-01 · Boyi Kang, Xinfa Zhu, Zihan Zhang, Zhen Ye 외

Recent advancements in language models (LMs) have demonstrated strong capabilities in semantic understanding and contextual modeling, which have flourished in generative speech enhancement (SE). However, many LM-based SE…

Language ModelingLanguage ModellingSpeech Enhancement

Ancient Greek to Modern Greek Machine Translation: A Novel Benchmark and Fine-Tuning Experiments on LLMs and NMT Models

2026-05-18 · Spyridon Mavromatis, Sokratis Sofianopoulos, Prokopis Prokopidis, Maria Giagkou arxiv

Machine Translation (MT) for Ancient Greek (AG) to Modern Greek (MG) is a low-resource task, constrained by the lack of large-scale, high-quality parallel data. We address this gap by introducing the AG-MG Parallel Corpu…

Machine Translation

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement

2025-10-23 · Haoyin Yan, Chengwei Liu, Shaofei Xue, Xiaotao Liang 외 arxiv

Neural audio codecs have largely promoted the application of language models (LMs) for speech applications. However, the effectiveness of autoregressive LM-based models in unifying speech enhancement (SE) tasks remains u…

Reinforcement LearningSpeech EnhancementSpeech Separation