RusTitW: Russian Language Text Dataset for Visual Text in-the-Wild Recognition
Information surrounds people in modern life. Text is a very efficient type of information that people use for communication for centuries. However, automated text-in-the-wild recognition remains a challenging problem. The major limitation for a DL system is the lack of training data. For the competitive performance, training set must contain many samples that replicate the real-world cases. While there are many high-quality datasets for English text recognition; there are no available datasets for Russian language. In this paper, we present a large-scale human-labeled dataset for Russian text recognition in-the-wild. We also publish a synthetic dataset and code to reproduce the generation process
Code (1)
Similar Papers 제목 키워드 기반
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
Text-to-image generation models have gained popularity among users around the world. However, many of these models exhibit a strong bias toward English-speaking cultures, ignoring or misrepresenting the unique characteri…
Image GenerationText to Image GenerationText-to-Image GenerationCreating an Aligned Russian Text Simplification Dataset from Language Learner Data
Parallel language corpora where regular texts are aligned with their simplified versions can be used in both natural language processing and theoretical linguistic studies. They are essential for the task of automatic te…
Text SimplificationPhonetic and Visual Priors for Decipherment of Informal Romanization
Informal romanization is an idiosyncratic process used by humans in informal digital communication to encode non-Latin script languages into Latin character sets found on common keyboards. Character substitution choices …
DeciphermentInductive BiasDataset for Automatic Summarization of Russian News
Automatic text summarization has been studied in a variety of domains and languages. However, this does not hold for the Russian language. To overcome this issue, we present Gazeta, the first dataset for summarization of…
Text SummarizationvalidLANGUAGE MODEL EMBEDDINGS IMPROVE SENTIMENT ANALYSIS IN RUSSIAN
Sentiment analysis is one of the most popular natural language processing tasks. In this paper we introduce pre-trained Russian language models which are used to extract embeddings (ELMo) to improve accuracy for classifi…
ArticlesLanguage ModelingLanguage Modellingmodel+2