paper-with-me

홈 › Papers

A Large Dataset of Historical Japanese Documents with Complex Layouts

2020-04-18 · Zejiang Shen, Kaixuan Zhang, Melissa Dell

Deep learning-based approaches for automatic document layout analysis and content extraction have the potential to unlock rich information trapped in historical documents on a large scale. One major hurdle is the lack of large datasets for training robust models. In particular, little training data exist for Asian languages. To this end, we present HJDataset, a Large Dataset of Historical Japanese Documents with Complex Layouts. It contains over 250,000 layout element annotations of seven types. In addition to bounding boxes and masks of the content regions, it also includes the hierarchical structures and reading orders for layout elements. The dataset is constructed using a combination of human and machine efforts. A semi-rule based method is developed to extract the layout elements, and the results are checked by human inspectors. The resulting large-scale dataset is used to provide baseline performance analyses for text region detection using state-of-the-art deep learning models. And we demonstrate the usefulness of the dataset on real-world document digitization tasks. The dataset is available at https://dell-research-harvard.github.io/HJDataset/.

📄 PDF Abstract BibTeX arXiv:2004.08686

Code (3)

dell-research-harvard/HJDataset 공식 구현 pytorch
Layout-Parser/layout-parser pytorch
felipeescallon/layout_parser pytorch

Tasks

Document Layout Analysis

Similar Papers 제목 키워드 기반

A human-inspired recognition system for premodern Japanese historical documents

2019-05-14 · Anh Duc Le, Tarin Clanuwat, Asanobu Kitamoto

Recognition of historical documents is a challenging problem due to the noised, damaged characters and background. However, in Japanese historical documents, not only contains the mentioned problems, pre-modern Japanese …

KuroNet: Pre-Modern Japanese Kuzushiji Character Recognition with Deep Learning

2019-10-21 · Tarin Clanuwat, Alex Lamb, Asanobu Kitamoto

Kuzushiji, a cursive writing style, had been used in Japan for over a thousand years starting from the 8th century. Over 3 millions books on a diverse array of topics, such as literature, science, mathematics and even co…

Deep Learning

Predicting the Ordering of Characters in Japanese Historical Documents

2021-06-12 · Alex Lamb, Tarin Clanuwat, Siyu Han, Mikel Bober-Irizar 외

Japan is a unique country with a distinct cultural heritage, which is reflected in billions of historical documents that have been preserved. However, the change in Japanese writing system in 1900 made these documents in…

Language ModelingLanguage ModellingMachine TranslationWord Embeddings

Automated Transcription for Pre-Modern Japanese Kuzushiji Documents by Random Lines Erasure and Curriculum Learning

2020-05-06 · Anh Duc Le

Recognizing the full-page of Japanese historical documents is a challenging problem due to the complex layout/background and difficulty of writing styles, such as cursive and connected characters. Most of the previous me…

Training Kindai OCR with parallel textline images and self-attention feature distance-based loss

2025-08-12 · Anh Le, Asanobu Kitamoto arxiv

Kindai documents, written in modern Japanese from the late 19th to early 20th century, hold significant historical value for researchers studying societal structures, daily life, and environmental conditions of that peri…

Domain Adaptation