paper-with-me

홈 › Papers

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space

2026-06-01 · Louis Mouchon arxiv

We present Echo, a proof-of-concept audio system built around a single 25 M-parameter ViT encoder. The encoder is pretrained with a JEPA objective and then specialised by stages to carry speaker identity, phonetic content, and dynamic source routing in the same 512-dimensional latent space, with no per-task fine-tuning at deployment. Light heads handle diarization (ArcFace + VBx) and dynamic source separation (null-target K-set prediction). On synthetic VoxCeleb2 mixtures with unknown K, the canonical stack reaches 15.00% blind DER, 97.80% PIT separation accuracy with +9.52 dB latent SI-SDR, and a +53.50-point speaker/content factorisation gap on a held-out k-NN probe. The point of Echo is not a new SOTA on any single task but the joint coexistence of three tasks on one encoder at this footprint. We document the design stage by stage, report the dead-ends, and identify the structural wall on end-to-end ASR through the VQ bottleneck that still bounds the PoC.

📄 PDF Abstract BibTeX arXiv:2606.01909

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker DiarizationSpeech Recognition

Similar Papers 제목 키워드 기반

A Conformer-based ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement and Speech Separation

2021-11-18 · Tom O'Malley, Arun Narayanan, Quan Wang, Alex Park 외

We present a frontend for improving robustness of automatic speech recognition (ASR), that jointly implements three modules within a single model: acoustic echo cancellation, speech enhancement, and speech separation. Th…

Acoustic echo cancellationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancement+3

NeuralEcho: A Self-Attentive Recurrent Neural Network For Unified Acoustic Echo Suppression And Speech Enhancement

2022-05-20 · Meng Yu, Yong Xu, Chunlei Zhang, Shi-Xiong Zhang 외

Acoustic echo cancellation (AEC) plays an important role in the full-duplex speech communication as well as the front-end speech enhancement for recognition in the conditions when the loudspeaker plays back. In this pape…

Acoustic echo cancellationSpeech Enhancementspeech-recognitionSpeech Recognition

A Universally-Deployable ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement, and Voice Separation

2022-09-14 · Tom O'Malley, Arun Narayanan, Quan Wang

Recent work has shown that it is possible to train a single model to perform joint acoustic echo cancellation (AEC), speech enhancement, and voice separation, thereby serving as a unified frontend for robust automatic sp…

Acoustic echo cancellationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancement+2

Real-Time Joint Personalized Speech Enhancement and Acoustic Echo Cancellation

2022-11-04 · Sefik Emre Eskimez, Takuya Yoshioka, Alex Ju, Min Tang 외

Personalized speech enhancement (PSE) is a real-time SE approach utilizing a speaker embedding of a target person to remove background noise, reverberation, and interfering voices. To deploy a PSE model for full duplex c…

Acoustic echo cancellationMulti-Task LearningSpeech Enhancement

End-To-End Deep Learning-based Adaptation Control for Linear Acoustic Echo Cancellation

2023-06-04 · Thomas Haubner, Andreas Brendel, Walter Kellermann

The attenuation of acoustic loudspeaker echoes remains to be one of the open challenges to achieve pleasant full-duplex hands free speech communication. In many modern signal enhancement interfaces, this problem is addre…

Acoustic echo cancellation