paper-with-me

홈 › Papers

EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs

2025-09-11 · Yuhao Zhang, Yuhao Du, Zhanchen Dai, Xiangnan Ma, Kaiqi Kou, Benyou Wang, Haizhou Li arxiv

Speech-to-speech large language models (SLLMs) are attracting increasing attention. Derived from text-based large language models (LLMs), SLLMs often exhibit degradation in knowledge and reasoning capabilities. We hypothesize that this limitation arises because current training paradigms for SLLMs fail to bridge the acoustic-semantic gap in the feature representation space. To address this issue, we propose EchoX, which leverages semantic representations and dynamically generates speech training targets. This approach integrates both acoustic and semantic learning, enabling EchoX to preserve strong reasoning abilities as a speech LLM. Experimental results demonstrate that EchoX, with about six thousand hours of training data, achieves advanced performance on multiple knowledge-based question-answering benchmarks. The project is available at https://github.com/FreedomIntelligence/EchoX.

📄 PDF Abstract BibTeX arXiv:2509.09174

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EchoXFlow: A Beamspace Echocardiography Dataset for Cardiac Motion, Flow, and Function

2026-05-06 · Elias Stenhede, Joanna Sulkowska, Eivind Bjørkan Orstad, Henrik Schirmer 외 arxiv

We introduce EchoXFlow, a clinical echocardiography dataset for learning from ultrasound in its native acquisition geometry rather than from scan-converted Cartesian videos. Existing public datasets offer limited opportu…

Residual acoustic echo suppression based on efficient multi-task convolutional neural network

2020-09-29 · Xinquan Zhou, Yanhong Leng

Acoustic echo degrades the user experience in voice communication systems thus needs to be suppressed completely. We propose a real-time residual acoustic echo suppression (RAES) method using an efficient convolutional n…

Multi-Task Learning

Deep Learning for Joint Acoustic Echo and Acoustic Howling Suppression in Hybrid Meetings

2023-05-02 · Hao Zhang, Meng Yu, Dong Yu

Hybrid meetings have become increasingly necessary during the post-COVID period and also brought new challenges for solving audio-related problems. In particular, the interplay between acoustic echo and acoustic howling …

Speech Separation

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs

2026-05-11 · Liang Lin, Chunxi Luo, Kaiwen Luo, Jie Zhang 외 arxiv

Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. Existing robustness methods primarily rely on waveform-level acoustic enhanceme…

Adaptive Speech Quality Aware Complex Neural Network for Acoustic Echo Cancellation with Supervised Contrastive Learning

2022-10-30 · Bozhong Liu, Xiaoxi Yu, Hantao Huang

Acoustic echo cancellation (AEC) is designed to remove echoes, reverberation, and unwanted added sounds from the microphone signal while maintaining the quality of the near-end speaker's speech. This paper proposes adapt…

Acoustic echo cancellationContrastive Learning