paper-with-me

Papers

DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech Translation

2023-10-26 · Yongxin Zhu, Zhujin Gao, Xinyuan Zhou, Zhongyi Ye, Linli Xu

While Diffusion Generative Models have achieved great success on image generation tasks, how to efficiently and effectively incorporate them into speech generation especially translation tasks remains a non-trivial problem. Specifically, due to the low information density of speech data, the transformed discrete speech unit sequence is much longer than the corresponding text transcription, posing significant challenges to existing auto-regressive models. Furthermore, it is not optimal to brutally apply discrete diffusion on the speech unit sequence while disregarding the continuous space structure, which will degrade the generation performance significantly. In this paper, we propose a novel diffusion model by applying the diffusion forward process in the \textit{continuous} speech representation space, while employing the diffusion backward process in the \textit{discrete} speech unit space. In this way, we preserve the semantic structure of the continuous speech representation space in the diffusion process and integrate the continuous and discrete diffusion models. We conduct extensive experiments on the textless direct speech-to-speech translation task, where the proposed method achieves comparable results to the computationally intensive auto-regressive baselines (500 steps on average) with significantly fewer decoding steps (50 steps).

📄 PDF Abstract BibTeX arXiv:2310.17570

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationSpeech-to-Speech TranslationTranslation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

DiffSLT: Enhancing Diversity in Sign Language Translation via Diffusion Model

2024-11-26 · Jihwan Moon, Jihoon Park, Jungeun Kim, Jongseong Bae 외

Sign language translation (SLT) is challenging, as it involves converting sign language videos into natural language. Previous studies have prioritized accuracy over diversity. However, diversity is crucial for handling …

DiversityMachine TranslationSign Language TranslationTranslation

TextDiffSeg: Text-guided Latent Diffusion Model for 3d Medical Images Segmentation

2025-04-16 · Kangbo Ma

Diffusion Probabilistic Models (DPMs) have demonstrated significant potential in 3D medical image segmentation tasks. However, their high computational cost and inability to fully capture global 3D contextual information…

Image SegmentationLatent Diffusion Model for 3DMedical Image SegmentationOrgan Segmentation+2

DiffSwap++: 3D Latent-Controlled Diffusion for Identity-Preserving Face Swapping

2025-11-04 · Weston Bondurant, Arkaprava Sinha, Hieu Le, Srijan Das 외 arxiv

Diffusion-based approaches have recently achieved strong results in face swapping, offering improved visual quality over traditional GAN-based methods. However, even state-of-the-art models often suffer from fine-grained…

Face Swapping

DiffStyler: Diffusion-based Localized Image Style Transfer

2024-03-27 · Shaoxu Li

Image style transfer aims to imbue digital imagery with the distinctive attributes of style targets, such as colors, brushstrokes, shapes, whilst concurrently preserving the semantic integrity of the content. Despite the…

DenoisingStyle Transfer

DIffSteISR: Harnessing Diffusion Prior for Superior Real-world Stereo Image Super-Resolution

2024-08-14 · Yuanbo Zhou, Xinlin Zhang, Wei Deng, Tao Wang 외

We introduce DiffSteISR, a pioneering framework for reconstructing real-world stereo images. DiffSteISR utilizes the powerful prior knowledge embedded in pre-trained text-to-image model to efficiently recover the lost te…

Image Super-ResolutionStereo Image Super-ResolutionSuper-ResolutionTAG