paper-with-me

홈 › Papers

DuoSearch: A Novel Search Engine for Bulgarian Historical Documents

2023-05-30 · Angel Beshirov, Suzan Hadzhieva, Ivan Koychev, Milena Dobreva

Search in collections of digitised historical documents is hindered by a two-prong problem, orthographic variety and optical character recognition (OCR) mistakes. We present a new search engine for historical documents, DuoSearch, which uses ElasticSearch and machine learning methods based on deep neural networks to offer a solution to this problem. It was tested on a collection of historical newspapers in Bulgarian from the mid-19th to the mid-20th century. The system provides an interactive and intuitive interface for the end-users allowing them to enter search terms in modern Bulgarian and search across historical spellings. This is the first solution facilitating the use of digitised historical documents in Bulgarian.

📄 PDF Abstract BibTeX arXiv:2305.19392

Code (1)

angelbeshirov/duosearch 공식 구현

Tasks

Optical Character RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Post-OCR Text Correction for Bulgarian Historical Documents

2024-08-31 · Angel Beshirov, Milena Dobreva, Dimitar Dimitrov, Momchil Hardalov 외

The digitization of historical documents is crucial for preserving the cultural heritage of the society. An important step in this process is converting scanned images to text using Optical Character Recognition (OCR), w…

Optical Character RecognitionOptical Character Recognition (OCR)

Categorisation of Bulgarian Legislative Documents

2020-09-01 · CLIB 2020 9 · Nikola Obreshkov, Martin Yalamov, Svetla Koeva

The paper presents the categorisation of Bulgarian MARCELL corpus in toplevel EuroVoc domains. The Bulgarian MARCELL corpus is part of a recently developed multilingual corpus representing the national legislation in sev…

Term Extraction

Natural Language Processing Pipeline to Annotate Bulgarian Legislative Documents

2020-05-01 · LREC 2020 5 · Svetla Koeva, Nikola Obreshkov, Martin Yalamov

The paper presents the Bulgarian MARCELL corpus, part of a recently developed multilingual corpus representing the national legislation in seven European countries and the NLP pipeline that turns the web crawled data int…

Sentence

CONDITOR1: Topic Maps and DITA labelling tool for textual documents with historical information

2016-03-23 · Piedad Garrido, Jesus Tramullas, Manuel Coll

Conditor is a software tool which works with textual documents containing historical information. The purpose of this work two-fold: firstly to show the validity of the developed engine to correctly identify and label th…

Information RetrievalRecommendation SystemsRetrieval

Syntactic and morphological features after verbs of perception: Bulgarian in Balkan context

2020-09-01 · CLIB 2020 9 · Ekaterina Tarpomanova

The paper analyses the types of constructions that express a subordinate event after a verb of perception in the languages of the Balkan Sprachbund. The subordinate clauses that may follow a verb of perception are a resu…