paper-with-me

홈 › Papers

Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems

2025-06-16 · Tuan Nguyen, Long-Vu Hoang, Huy-Dat Tran

This paper presents our system for the MLC-SLM Challenge 2025, focusing on multilingual speech recognition and language modeling with large language models (LLMs). Our approach combines a fine-tuned Whisper-large-v3 encoder with efficient projector architectures and various decoder configurations. We employ a three-stage training methodology that progressively optimizes the encoder, projector, and LLM components. Our system achieves competitive performance with a private test average WER/CER result of 16.63% using the Gemma3-12B and 18.6% using the Qwen2.5-7B as decoder-only language model.

📄 PDF Abstract BibTeX arXiv:2506.13596

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderLanguage ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario

2024-12-01 · Zheshu Song, Ziyang Ma, Yifan Yang, Jianheng Zhuo 외

Large Language Models (LLMs) have showcased exceptional performance across diverse NLP tasks, and their integration with speech encoder is rapidly emerging as a dominant trend in the Automatic Speech Recognition (ASR) fi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Geometry of Ordinal Representations in Language Models

2026-07-05 · Saksham Bassi, Sharvi Tomar arxiv

Recent work showed that language models represent character counts on curved 1D manifolds, with attention heads performing geometric transformations to enable computation. We test whether this generalizes across four ord…

RIPA: Sensory-Vector Prompt Injection Attacks on LLM-Controlled ROS 2 Robots

2026-06-26 · Nima Dorzhiev arxiv

We present RIPA, the first systematic multi-channel empirical study of prompt injection attacks delivered through the sensory pipeline of a ROS 2-based LLM-controlled robotic system. Across 100 independent runs per injec…

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

2026-06-25 · Subham Kumar, Prakrithi Shivaprakash, Abhishek Manoharan, Astut Kurariya 외 arxiv

Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown. In this study, w…

Speech Recognition

Multimodal Models Meet Presentation Attack Detection on ID Documents

2026-03-31 · Marina Villanueva, Juan M. Espin, Juan E. Tapia arxiv

The integration of multimodal models into Presentation Attack Detection (PAD) for ID Documents represents a significant advancement in biometric security. Traditional PAD systems rely solely on visual features, which oft…