paper-with-me

홈 › Papers

RusTitW: Russian Language Text Dataset for Visual Text in-the-Wild Recognition

2023-03-29 · Igor Markov, Sergey Nesteruk, Andrey Kuznetsov, Denis Dimitrov

Information surrounds people in modern life. Text is a very efficient type of information that people use for communication for centuries. However, automated text-in-the-wild recognition remains a challenging problem. The major limitation for a DL system is the lack of training data. For the competitive performance, training set must contain many samples that replicate the real-world cases. While there are many high-quality datasets for English text recognition; there are no available datasets for Russian language. In this paper, we present a large-scale human-labeled dataset for Russian text recognition in-the-wild. We also publish a synthetic dataset and code to reproduce the generation process

📄 PDF Abstract BibTeX arXiv:2303.16531

Code (1)

markovivl/synthtext 공식 구현

Similar Papers 제목 키워드 기반

RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation

2025-02-11 · Viacheslav Vasilev, Julia Agafonova, Nikolai Gerasimenko, Alexander Kapitanov 외

Text-to-image generation models have gained popularity among users around the world. However, many of these models exhibit a strong bias toward English-speaking cultures, ignoring or misrepresenting the unique characteri…

Image GenerationText to Image GenerationText-to-Image Generation

Creating an Aligned Russian Text Simplification Dataset from Language Learner Data

2021-04-01 · EACL (BSNLP) 2021 4 · Anna Dmitrieva, Jörg Tiedemann

Parallel language corpora where regular texts are aligned with their simplified versions can be used in both natural language processing and theoretical linguistic studies. They are essential for the task of automatic te…

Text Simplification

Phonetic and Visual Priors for Decipherment of Informal Romanization

2020-05-05 · ACL 2020 6 · Maria Ryskina, Matthew R. Gormley, Taylor Berg-Kirkpatrick

Informal romanization is an idiosyncratic process used by humans in informal digital communication to encode non-Latin script languages into Latin character sets found on common keyboards. Character substitution choices …

DeciphermentInductive Bias

Dataset for Automatic Summarization of Russian News

2020-06-19 · Ilya Gusev

Automatic text summarization has been studied in a variety of domains and languages. However, this does not hold for the Russian language. To overcome this issue, we present Gazeta, the first dataset for summarization of…

Text Summarizationvalid

LANGUAGE MODEL EMBEDDINGS IMPROVE SENTIMENT ANALYSIS IN RUSSIAN

2019-05-29 · Computational Linguistics and Intellectual Technologies: Proceedings of the International Conference “Dialogue 2019” 2019 5 · Baymurzina D. R., Kuznetsov D. P., Burtsev M. S.

Sentiment analysis is one of the most popular natural language processing tasks. In this paper we introduce pre-trained Russian language models which are used to extract embeddings (ELMo) to improve accuracy for classifi…

ArticlesLanguage ModelingLanguage Modellingmodel+2