paper-with-me

Papers

Leveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translation

2022-03-24 · Ye Jia, Yifan Ding, Ankur Bapna, Colin Cherry, Yu Zhang, Alexis Conneau, Nobuyuki Morioka

End-to-end speech-to-speech translation (S2ST) without relying on intermediate text representations is a rapidly emerging frontier of research. Recent works have demonstrated that the performance of such direct S2ST systems is approaching that of conventional cascade S2ST when trained on comparable datasets. However, in practice, the performance of direct S2ST is bounded by the availability of paired S2ST training data. In this work, we explore multiple approaches for leveraging much more widely available unsupervised and weakly-supervised speech and text data to improve the performance of direct S2ST based on Translatotron 2. With our most effective approaches, the average translation quality of direct S2ST on 21 language pairs on the CVSS-C corpus is improved by +13.6 BLEU (or +113% relatively), as compared to the previous state-of-the-art trained without additional data. The improvements on low-resource language are even more significant (+398% relatively on average). Our comparative studies suggest future research directions for S2ST and speech representation learning.

📄 PDF Abstract BibTeX arXiv:2203.13339

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSpeech Representation LearningSpeech-to-Speech TranslationTranslation

Similar Papers 제목 키워드 기반

MAF: Multimodal Alignment Framework for Weakly-Supervised Phrase Grounding

2020-10-12 · EMNLP 2020 11 · Qinxin Wang, Hao Tan, Sheng Shen, Michael W. Mahoney 외

Phrase localization is a task that studies the mapping from textual phrases to regions of an image. Given difficulties in annotating phrase-to-object datasets at scale, we develop a Multimodal Alignment Framework (MAF) t…

Phrase Grounding

Unsupervised Deep Tracking

2019-04-03 · CVPR 2019 6 · Ning Wang, Yibing Song, Chao Ma, Wengang Zhou 외

We propose an unsupervised visual tracking method in this paper. Different from existing approaches using extensive annotated data for supervised learning, our CNN model is trained on large-scale unlabeled videos in an u…

Visual Tracking

AnoFPDM: Anomaly Segmentation with Forward Process of Diffusion Models for Brain MRI

2024-04-24 · Yiming Che, Fazle Rafsani, Jay Shah, Md Mahfuzur Rahman Siddiquee 외

Weakly-supervised diffusion models (DMs) in anomaly segmentation, leveraging image-level labels, have attracted significant attention for their superior performance compared to unsupervised methods. It eliminates the nee…

Anomaly SegmentationSegmentation

Improving Weakly-supervised Video Instance Segmentation by Leveraging Spatio-temporal Consistency

2024-08-29 · Farnoosh Arefi, Amir M. Mansourian, Shohreh Kasaei

The performance of Video Instance Segmentation (VIS) methods has improved significantly with the advent of transformer networks. However, these networks often face challenges in training due to the high annotation cost. …

Instance SegmentationSegmentationSemantic SegmentationVideo Instance Segmentation

PatchMVSNet: Patch-wise Unsupervised Multi-View Stereo for Weakly-Textured Surface Reconstruction

2022-03-04 · Haonan Dong, Jian Yao

Learning-based multi-view stereo (MVS) has gained fine reconstructions on popular datasets. However, supervised learning methods require ground truth for training, which is hard to be collected, especially for the large-…

Depth EstimationSurface Reconstruction