DavarOCR: A Toolbox for OCR and Multi-Modal Document Understanding
This paper presents DavarOCR, an open-source toolbox for OCR and document understanding tasks. DavarOCR currently implements 19 advanced algorithms, covering 9 different task forms. DavarOCR provides detailed usage instructions and the trained models for each algorithm. Compared with the previous opensource OCR toolbox, DavarOCR has relatively more complete support for the sub-tasks of the cutting-edge technology of document understanding. In order to promote the development and application of OCR technology in academia and industry, we pay more attention to the use of modules that different sub-domains of technology can share. DavarOCR is publicly released at https://github.com/hikopensource/Davar-Lab-OCR.
Code (1)
Tasks
document understandingOptical Character Recognition (OCR)Similar Papers 제목 키워드 기반
MMRec: Simplifying Multimodal Recommendation
This paper presents an open-source toolbox, MMRec for multimodal recommendation. MMRec simplifies and canonicalizes the process of implementing and comparing multimodal recommendation models. The objective of MMRec is to…
Multimodal RecommendationPROSO Toolbox: a unified protein-constrained genome-scale modelling framework for strain designing and optimization
The genome-scale metabolic model with protein constraint (PC-model) has been increasingly popular for microbial metabolic simulations. We present PROSO Toolbox, a unified and simple-to-use PC-model toolbox that takes any…
MultiCalib4DEB: A toolbox exploiting multimodal optimisation in Dynamic Energy Budget parameters calibration
Calibration is a crucial step for the validation of computational models and a challenging task to accomplish. Dynamic Energy Budget (DEB) theory has experienced an exponential rise in the number of published papers, whi…
BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
Early children's developmental trajectories set up a natural goal for sample-efficient pretraining of vision foundation models. We introduce BabyVLM-V2, a developmentally grounded framework for infant-inspired vision-lan…
Spatial ReasoningDocopilot: Improving Multimodal Models for Document-Level Understanding
Despite significant progress in multimodal large language models (MLLMs), their performance on complex, multi-page document comprehension remains inadequate, largely due to the lack of high-quality, document-level datase…