paper-with-me

홈 › Papers

End-to-End Subtitle Detection and Recognition for Videos in East Asian Languages via CNN Ensemble with Near-Human-Level Performance

2016-11-18 · Yan Xu, Siyuan Shan, Ziming Qiu, Zhipeng Jia, Zhengyang Shen, Yipei Wang, Mengfei Shi, Eric I-Chao Chang

In this paper, we propose an innovative end-to-end subtitle detection and recognition system for videos in East Asian languages. Our end-to-end system consists of multiple stages. Subtitles are firstly detected by a novel image operator based on the sequence information of consecutive video frames. Then, an ensemble of Convolutional Neural Networks (CNNs) trained on synthetic data is adopted for detecting and recognizing East Asian characters. Finally, a dynamic programming approach leveraging language models is applied to constitute results of the entire body of text lines. The proposed system achieves average end-to-end accuracies of 98.2% and 98.3% on 40 videos in Simplified Chinese and 40 videos in Traditional Chinese respectively, which is a significant outperformance of other existing methods. The near-perfect accuracy of our system dramatically narrows the gap between human cognitive ability and state-of-the-art algorithms used for such a task.

📄 PDF Abstract BibTeX arXiv:1611.06159

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

NLP Driven Ensemble Based Automatic Subtitle Generation and Semantic Video Summarization Technique

2019-04-22 · VB Aswin, Mohammed Javed, Parag Parihar, K Aswanth 외

This paper proposes an automatic subtitle generation and semantic video summarization technique. The importance of automatic video summarization is vast in the present era of big data. Video summarization helps in effici…

speech-recognitionSpeech RecognitionText SummarizationVideo Summarization

Multi-language Video Subtitle Dataset for Image-based Text Recognition

2024-11-07 · Thanadol Singkhornart, Olarik Surinta

The Multi-language Video Subtitle Dataset is a comprehensive collection designed to support research in text recognition across multiple languages. This dataset includes 4,224 subtitle images extracted from 24 videos sou…

Computational Efficiency

Aligning Subtitles in Sign Language Videos

2021-05-06 · ICCV 2021 10 · Hannah Bull, Triantafyllos Afouras, Gül Varol, Samuel Albanie 외

The goal of this work is to temporally align asynchronous subtitles in sign language videos. In particular, we focus on sign-language interpreted TV broadcast data comprising (i) a video of continuous signing, and (ii) s…

Machine TranslationTranslation

Cloud-based Automatic Speech Recognition Systems for Southeast Asian Languages

2022-10-07 · Lei Wang, Rong Tong, Cheung Chi Leung, Sunil Sivadas 외

This paper provides an overall introduction of our Automatic Speech Recognition (ASR) systems for Southeast Asian languages. As not much existing work has been carried out on such regional languages, a few difficulties s…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

HowToCaption: Prompting LLMs to Transform Video Annotations at Scale

2023-10-07 · Nina Shvetsova, Anna Kukleva, Xudong Hong, Christian Rupprecht 외

Instructional videos are a common source for learning text-video or even multimodal representations by leveraging subtitles extracted with automatic speech recognition systems (ASR) from the audio signal in the videos. H…

Automatic Speech RecognitionVideo CaptioningVideo RetrievalZero-Shot Video-Audio Retrieval+1