paper-with-me

홈 › Papers

Tutorial Proposal: End-to-End Speech Translation

2021-04-01 · EACL 2021 2 · Jan Niehues, Elizabeth Salesky, Marco Turchi, Matteo Negri

Speech translation is the translation of speech in one language typically to text in another, traditionally accomplished through a combination of automatic speech recognition and machine translation. Speech translation has attracted interest for many years, but the recent successful applications of deep learning to both individual tasks have enabled new opportunities through joint modeling, in what we today call {`}end-to-end speech translation.{'} In this tutorial we will introduce the techniques used in cutting-edge research on speech translation. Starting from the traditional cascaded approach, we will given an overview on data sources and model architectures to achieve state-of-the art performance with end-to-end speech translation for both high- and low-resource languages. In addition, we will discuss methods to evaluate analyze the proposed solutions, as well as the challenges faced when applying speech translation models for real-world applications.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognitionSpeech RecognitionTranslation

Similar Papers 제목 키워드 기반

Simultaneous Translation

2020-11-01 · EMNLP 2020 11 · Liang Huang, Colin Cherry, Mingbo Ma, Naveen Arivazhagan 외

Simultaneous translation, which performs translation concurrently with the source speech, is widely useful in many scenarios such as international conferences, negotiations, press releases, legal proceedings, and medicin…

Machine Translationspeech-recognitionSpeech RecognitionSpeech Synthesis+1

Long-form Simultaneous Speech Translation: Thesis Proposal

2023-10-17 · Peter Polák

Simultaneous speech translation (SST) aims to provide real-time translation of spoken language, even before the speaker finishes their sentence. Traditionally, SST has been addressed primarily by cascaded systems that de…

FormMachine TranslationSegmentationSentence+3

JoeyS2T: Minimalistic Speech-to-Text Modeling with JoeyNMT

2022-10-05 · Mayumi Ohta, Julia Kreutzer, Stefan Riezler

JoeyS2T is a JoeyNMT extension for speech-to-text tasks such as automatic speech recognition and end-to-end speech translation. It inherits the core philosophy of JoeyNMT, a minimalist NMT toolkit built on PyTorch, seeki…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderNMT+5

Multimodal Machine Learning: Integrating Language, Vision and Speech

2017-07-01 · ACL 2017 7 · Louis-Philippe Morency, Tadas Baltru{\v{s}}aitis

Multimodal machine learning is a vibrant multi-disciplinary research field which addresses some of the original goals of artificial intelligence by integrating and modeling multiple communicative modalities, including li…

Audio-Visual Speech RecognitionBIG-bench Machine LearningImage CaptioningQuestion Answering+8

Data Expansion using Back Translation and Paraphrasing for Hate Speech Detection

2021-05-25 · Djamila Romaissa Beddiar, Md Saroar Jahan, Mourad Oussalah

With proliferation of user generated contents in social media platforms, establishing mechanisms to automatically identify toxic and abusive content becomes a prime concern for regulators, researchers, and society. Keepi…

Data AugmentationDecoderDeep LearningHate Speech Detection+3