OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles
We present a new major release of the OpenSubtitles collection of parallel corpora. The release is compiled from a large database of movie and TV subtitles and includes a total of 1689 bitexts spanning 2.6 billion sentences across 60 languages. The release also incorporates a number of enhancements in the preprocessing and alignment of the subtitles, such as the automatic correction of OCR errors and the use of meta-data to estimate the quality of each subtitle and score subtitle pairs.
Code (0)
등록된 구현이 없습니다.
Tasks
Optical Character Recognition (OCR)Similar Papers 제목 키워드 기반
OpenSubtitles2018: Statistical Rescoring of Sentence Alignments in Large, Noisy Parallel Corpora
Fine-grained Emotion and Intent Learning in Movie Dialogues
We propose a novel large-scale emotional dialogue dataset, consisting of 1M dialogues retrieved from the OpenSubtitles corpus and annotated with 32 emotions and 9 empathetic response intents using a BERT-based fine-grain…
Evaluating Gender Bias Transfer from Film Data
Films are a rich source of data for natural language processing. OpenSubtitles (Lison and Tiedemann, 2016) is a popular movie script dataset, used for training models for tasks such as machine translation and dialogue ge…
Dialogue GenerationMachine TranslationSentenceSentence Embedding+2Innovations in Parallel Corpus Search Tools
Recent years have seen an increased interest in and availability of parallel corpora. Large corpora from international organizations (e.g. European Union, United Nations, European Patent Office), or from multilingual Int…
Machine TranslationSentenceTranslationWord Alignment+1Building a Hebrew Semantic Role Labeling Lexical Resource from Parallel Movie Subtitles
We present a semantic role labeling resource for Hebrew built semi-automatically through annotation projection from English. This corpus is derived from the multilingual OpenSubtitles dataset and includes short informal …
Morphological AnalysisSemantic Role Labeling