paper-with-me

Papers

Bi-Text Alignment of Movie Subtitles for Spoken English-Arabic Statistical Machine Translation

2016-09-05 · Fahad Al-Obaidli, Stephen Cox, Preslav Nakov

We describe efforts towards getting better resources for English-Arabic machine translation of spoken text. In particular, we look at movie subtitles as a unique, rich resource, as subtitles in one language often get translated into other languages. Movie subtitles are not new as a resource and have been explored in previous research; however, here we create a much larger bi-text (the biggest to date), and we further generate better quality alignment for it. Given the subtitles for the same movie in different languages, a key problem is how to align them at the fragment level. Typically, this is done using length-based alignment, but for movie subtitles, there is also time information. Here we exploit this information to develop an original algorithm that outperforms the current best subtitle alignment tool, subalign. The evaluation results show that adding our bi-text to the IWSLT training bi-text yields an improvement of over two BLEU points absolute.

📄 PDF Abstract BibTeX arXiv:1609.01188

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Dual Subtitles as Parallel Corpora

2014-05-01 · LREC 2014 5 · Shikun Zhang, Wang Ling, Chris Dyer

In this paper, we leverage the existence of dual subtitles as a source of parallel data. Dual subtitles present viewers with two languages simultaneously, and are generally aligned in the segment level, which removes the…

Machine TranslationSentenceTranslationWord Sense Disambiguation

OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles

2016-05-01 · LREC 2016 5 · Pierre Lison, J{\"o}rg Tiedemann

We present a new major release of the OpenSubtitles collection of parallel corpora. The release is compiled from a large database of movie and TV subtitles and includes a total of 1689 bitexts spanning 2.6 billion senten…

Optical Character Recognition (OCR)

StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles

2026-02-25 · Daniel Oliveira, David Martins de Matos arxiv

Visual storytelling models that correctly ground entities in images may still hallucinate semantic relationships, generating incorrect dialogue attribution, character interactions, or emotional states. We introduce Story…

Visual StorytellingVisual Grounding

Movie Question Answering: Remembering the Textual Cues for Layered Visual Contents

2018-04-25 · Bo Wang, Youjiang Xu, Yahong Han, Richang Hong

Movies provide us with a mass of visual content as well as attracting stories. Existing methods have illustrated that understanding movie stories through only visual content is still a hard problem. In this paper, for an…

Question AnsweringVideo Question Answering

Towards Visually-Guided Movie Subtitle Translation for Indic Languages

2026-05-12 · Tarun Chintada, Kshetrimayum Boynao Singh, Asif Ekbal arxiv

Movie subtitle translation is inherently multimodal, yet text-only systems often miss visual cues needed to convey emotion, action, and social nuance, especially for low-resource Indic languages (English to Hindi, Bengal…

Visual Grounding