DISGO: Automatic End-to-End Evaluation for Scene Text OCR
This paper discusses the challenges of optical character recognition (OCR) on natural scenes, which is harder than OCR on documents due to the wild content and various image backgrounds. We propose to uniformly use word error rates (WER) as a new measurement for evaluating scene-text OCR, both end-to-end (e2e) performance and individual system component performances. Particularly for the e2e metric, we name it DISGO WER as it considers Deletion, Insertion, Substitution, and Grouping/Ordering errors. Finally we propose to utilize the concept of super blocks to automatically compute BLEU scores for e2e OCR machine translation. The small SCUT public test set is used to demonstrate WER performance by a modularized OCR system.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationOptical Character RecognitionOptical Character Recognition (OCR)TranslationSimilar Papers 제목 키워드 기반
AI Model Disgorgement: Methods and Choices
Responsible use of data is an indispensable part of any machine learning (ML) implementation. ML developers must carefully collect and curate their datasets, and document their provenance. They must also make sure to res…
modelAutomatic Generation of German Drama Texts Using Fine Tuned GPT-2 Models
This study is devoted to the automatic generation of German drama texts. We suggest an approach consisting of two key steps: fine-tuning a GPT-2 model (the outline model) to generate outlines of scenes based on keywords …
SumTitles: a Summarization Dataset with Low Extractiveness
The existing dialogue summarization corpora are significantly extractive. We introduce a methodology for dataset extractiveness evaluation and present a new low-extractive corpus of movie dialogues for abstractive text s…
Abstractive Text SummarizationText SummarizationMCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark
The comprehension of text-rich visual scenes has become a focal point for evaluating Multi-modal Large Language Models (MLLMs) due to their widespread applications. Current benchmarks tailored to the scenario emphasize p…
FairnessScene Text RecognitionVisual ReasoningAutomatic Funny Scene Extraction from Long-form Cinematic Videos
Automatically extracting engaging and high-quality humorous scenes from cinematic titles is pivotal for creating captivating video previews and snackable content, boosting user engagement on streaming platforms. Long-for…
Scene Segmentation