paper-with-me

Papers

Callico: a Versatile Open-Source Document Image Annotation Platform

2024-05-02 · Christopher Kermorvant, Eva Bardou, Manon Blanco, Bastien Abadie

This paper presents Callico, a web-based open source platform designed to simplify the annotation process in document recognition projects. The move towards data-centric AI in machine learning and deep learning underscores the importance of high-quality data, and the need for specialised tools that increase the efficiency and effectiveness of generating such data. For document image annotation, Callico offers dual-display annotation for digitised documents, enabling simultaneous visualisation and annotation of scanned images and text. This capability is critical for OCR and HTR model training, document layout analysis, named entity recognition, form-based key value annotation or hierarchical structure annotation with element grouping. The platform supports collaborative annotation with versatile features backed by a commitment to open source development, high-quality code standards and easy deployment via Docker. Illustrative use cases - including the transcription of the Belfort municipal registers, the indexing of French World War II prisoners for the ICRC, and the extraction of personal information from the Socface project's census lists - demonstrate Callico's applicability and utility.

📄 PDF Abstract BibTeX arXiv:2405.01071

Code (0)

등록된 구현이 없습니다.

Tasks

Document Layout AnalysisHTRnamed-entity-recognitionNamed Entity RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

DOGE: Towards Versatile Visual Document Grounding and Referring

2024-11-26 · Yinan Zhou, Yuxin Chen, Haokun Lin, Shuyu Yang 외

In recent years, Multimodal Large Language Models (MLLMs) have increasingly emphasized grounding and referring capabilities to achieve detailed understanding and flexible user interaction. However, in the realm of visual…

document understanding

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding

2025-10-30 · Yiqiao Jin, Rachneet Kaur, Zhen Zeng, Sumitra Ganesh 외 arxiv

Multi-page visual documents such as manuals, brochures, presentations, and posters convey key information through layout, colors, icons, and cross-slide references. While multimodal large language models (MLLMs) offer op…

SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding

2024-08-27 · Chuanghao Ding, Xuejing Liu, Wei Tang, Juan Li 외

This paper introduces SynthDoc, a novel synthetic document generation pipeline designed to enhance Visual Document Understanding (VDU) by generating high-quality, diverse datasets that include text, images, tables, and c…

document understanding

jinns: a JAX Library for Physics-Informed Neural Networks

2024-12-18 · Hugo Gangloff, Nicolas Jouvin

jinns is an open-source Python library for physics-informed neural networks, built to tackle both forward and inverse problems, as well as meta-model learning. Rooted in the JAX ecosystem, it provides a versatile framewo…

PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents

2024-06-20 · Junjie Wang, Yin Zhang, Yatai Ji, Yuxiang Zhang 외

Recent advancements in Large Multimodal Models (LMMs) have leveraged extensive multimodal datasets to enhance capabilities in complex knowledge-driven tasks. However, persistent challenges in perceptual and reasoning err…