Processing M.A. Castrén's Materials: Multilingual Typed and Handwritten Manuscripts
The study forms a technical report of various tasks that have been performed on the materials collected and published by Finnish ethnographer and linguist, Matthias Alexander Castr\'en (1813-1852). The Finno-Ugrian Society is publishing Castr\'en's manuscripts as new critical and digital editions, and at the same time different research groups have also paid attention to these materials. We discuss the workflows and technical infrastructure used, and consider how datasets that benefit different computational tasks could be created to further improve the usability of these materials, and also to aid the further processing of similar archived collections. We specifically focus on the parts of the collections that are processed in a way that improves their usability in more technical applications, complementing the earlier work on the cultural and linguistic aspects of these materials. Most of these datasets are openly available in Zenodo. The study points to specific areas where further research is needed, and provides benchmarks for text recognition tasks.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Processing M.A. Castrén’s Materials: Multilingual Historical Typed and Handwritten Manuscripts
The study forms a technical report of various tasks that have been performed on the materials collected and published by Finnish ethnographer and linguist, Matthias Alexander Castrén (1813–1852). The Finno-Ugrian Society…
Surface Realization Shared Task 2018 (SR18): The Tilburg University Approach
This study describes the approach developed by the Tilburg University team to the shallow task of the Multilingual Surface Realization Shared Task 2018 (SR18). Based on (Castro Ferreira et al., 2017), the approach works …
Machine TranslationTranslationAdvancing Multilingual Handwritten Numeral Recognition with Attention-driven Transfer Learning
As deep learning continues to evolve, we have observed huge breakthroughs in the fields of medical imaging, video and frame generation, optical character recognition (OCR), and other domains. In the field of data analysi…
Handwritten Digit RecognitionOptical Character RecognitionOptical Character Recognition (OCR)Transfer LearningHW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
The proliferation of MultiLingual Visual Question Answering (MLVQA) benchmarks augments the capabilities of large language models (LLMs) and multi-modal LLMs, thereby enabling them to adeptly capture the intricate lingui…
Visual Question AnsweringHandwritten Script Identification from Text Lines
In a multilingual country like India where 12 different official scripts are in use, automatic identification of handwritten script facilitates many important applications such as automatic transcription of multilingual …
Optical Character RecognitionOptical Character Recognition (OCR)