paper-with-me

홈 › Papers

Cascaded Cross-Module Residual Learning towards Lightweight End-to-End Speech Coding

2019-06-18 · Kai Zhen, Jongmo Sung, Mi Suk Lee, Seung-Kwon Beack, Minje Kim

Speech codecs learn compact representations of speech signals to facilitate data transmission. Many recent deep neural network (DNN) based end-to-end speech codecs achieve low bitrates and high perceptual quality at the cost of model complexity. We propose a cross-module residual learning (CMRL) pipeline as a module carrier with each module reconstructing the residual from its preceding modules. CMRL differs from other DNN-based speech codecs, in that rather than modeling speech compression problem in a single large neural network, it optimizes a series of less-complicated modules in a two-phase training scheme. The proposed method shows better objective performance than AMR-WB and the state-of-the-art DNN-based speech codec with a similar network architecture. As an end-to-end model, it takes raw PCM signals as an input, but is also compatible with linear predictive coding (LPC), showing better subjective quality at high bitrates than AMR-WB and OPUS. The gain is achieved by using only 0.9 million trainable parameters, a significantly less complex architecture than the other DNN-based codecs in the literature.

📄 PDF Abstract BibTeX arXiv:1906.07769

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AuRA: Internalizing Audio Understanding into LLMs as LoRA

2026-06-09 · Bo Cheng, Lei Shi, Zhanyu Ma, Yuan Wu 외 arxiv

Recent efforts to extend large language models (LLMs) to speech inputs typically rely on cascaded ASR-LLM pipelines, end-to-end speech-language models, or bridge/distillation-based adaptation. While these routes respecti…

SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data

2024-02-10 · Hsuan-Fu Wang, Yi-Jen Shih, Heng-Jui Chang, Layne Berry 외

The recently proposed visually grounded speech model SpeechCLIP is an innovative framework that bridges speech and text through images via CLIP without relying on text transcription. On this basis, this paper introduces …

Keyword ExtractionMulti-Task LearningRepresentation LearningRetrieval

KIT’s IWSLT 2021 Offline Speech Translation System

2021-08-01 · ACL (IWSLT) 2021 8 · Tuan Nam Nguyen, Thai Son Nguyen, Christian Huber, Ngoc-Quan Pham 외

This paper describes KIT’submission to the IWSLT 2021 Offline Speech Translation Task. We describe a system in both cascaded condition and end-to-end condition. In the cascaded condition, we investigated different end-to…

Machine Translationspeech-recognitionSpeech RecognitionText Segmentation+1

DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors

2023-12-07 · Federico Landini, Mireia Diez, Themos Stafylakis, Lukáš Burget

Until recently, the field of speaker diarization was dominated by cascaded systems. Due to their limitations, mainly regarding overlapped speech and cumbersome pipelines, end-to-end models have gained great popularity la…

Decoderspeaker-diarizationSpeaker Diarization

Cascaded Residual Density Network for Crowd Counting

2021-07-29 · Kun Zhao, Luchuan Song, Bin Liu, Qi Chu 외

Crowd counting is a challenging task due to the issues such as scale variation and perspective variation in real crowd scenes. In this paper, we propose a novel Cascaded Residual Density Network (CRDNet) in a coarse-to-f…

Crowd Counting