paper-with-me

Papers

OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles

2016-05-01 · LREC 2016 5 · Pierre Lison, J{\"o}rg Tiedemann

We present a new major release of the OpenSubtitles collection of parallel corpora. The release is compiled from a large database of movie and TV subtitles and includes a total of 1689 bitexts spanning 2.6 billion sentences across 60 languages. The release also incorporates a number of enhancements in the preprocessing and alignment of the subtitles, such as the automatic correction of OCR errors and the use of meta-data to estimate the quality of each subtitle and score subtitle pairs.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

OpenSubtitles2018: Statistical Rescoring of Sentence Alignments in Large, Noisy Parallel Corpora

2018-05-01 · LREC 2018 5 · Pierre Lison, J{\"o}rg Tiedemann, Milen Kouylekov
Machine TranslationSentence

Fine-grained Emotion and Intent Learning in Movie Dialogues

2020-12-25 · Anuradha Welivita, Yubo Xie, Pearl Pu

We propose a novel large-scale emotional dialogue dataset, consisting of 1M dialogues retrieved from the OpenSubtitles corpus and annotated with 32 emotions and 9 empathetic response intents using a BERT-based fine-grain…

Evaluating Gender Bias Transfer from Film Data

2022-07-01 · NAACL (GeBNLP) 2022 7 · Amanda Bertsch, Ashley Oh, Sanika Natu, Swetha Gangu 외

Films are a rich source of data for natural language processing. OpenSubtitles (Lison and Tiedemann, 2016) is a popular movie script dataset, used for training models for tasks such as machine translation and dialogue ge…

Dialogue GenerationMachine TranslationSentenceSentence Embedding+2

Innovations in Parallel Corpus Search Tools

2014-05-01 · LREC 2014 5 · Martin Volk, Johannes Gra{\"e}n, Elena Callegaro

Recent years have seen an increased interest in and availability of parallel corpora. Large corpora from international organizations (e.g. European Union, United Nations, European Patent Office), or from multilingual Int…

Machine TranslationSentenceTranslationWord Alignment+1

Building a Hebrew Semantic Role Labeling Lexical Resource from Parallel Movie Subtitles

2020-05-17 · LREC 2020 5 · Ben Eyal, Michael Elhadad

We present a semantic role labeling resource for Hebrew built semi-automatically through annotation projection from English. This corpus is derived from the multilingual OpenSubtitles dataset and includes short informal …

Morphological AnalysisSemantic Role Labeling