paper-with-me

홈 › Papers

Hespi: A pipeline for automatically detecting information from hebarium specimen sheets

2024-10-11 · Robert Turnbull, Emily Fitzgerald, Karen Thompson, Joanne L. Birch

Specimen associated biodiversity data are sought after for biological, environmental, climate, and conservation sciences. A rate shift is required for the extraction of data from specimen images to eliminate the bottleneck that the reliance on human-mediated transcription of these data represents. We applied advanced computer vision techniques to develop the `Hespi' (HErbarium Specimen sheet PIpeline), which extracts a pre-catalogue subset of collection data on the institutional labels on herbarium specimens from their digital images. The pipeline integrates two object detection models; the first detects bounding boxes around text-based labels and the second detects bounding boxes around text-based data fields on the primary institutional label. The pipeline classifies text-based institutional labels as printed, typed, handwritten, or a combination and applies Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) for data extraction. The recognized text is then corrected against authoritative databases of taxon names. The extracted text is also corrected with the aide of a multimodal Large Language Model (LLM). Hespi accurately detects and extracts text for test datasets including specimen sheet images from international herbaria. The components of the pipeline are modular and users can train their own models with their own data and use them in place of the models provided.

📄 PDF Abstract BibTeX arXiv:2410.08740

Code (1)

rbturnbull/hespi 공식 구현

Tasks

Handwritten Text RecognitionHTRLanguage ModellingLarge Language ModelMultimodal Large Language Modelobject-detectionObject DetectionOptical Character RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Thespian: Multi-Character Text Role-Playing Game Agents

2023-08-03 · Christopher Cui, Xiangyu Peng, Mark Riedl

Text-adventure games and text role-playing games are grand challenges for reinforcement learning game playing agents. Text role-playing games are open-ended environments where an agent must faithfully play a particular c…

Few-Shot Learning

Automatically Identifying Comparator Groups on Twitter for Digital Epidemiology of Pregnancy Outcomes

2019-08-16 · Ari Z. Klein, Abeselom Gebreyesus, Graciela Gonzalez-Hernandez

Despite the prevalence of adverse pregnancy outcomes such as miscarriage, stillbirth, birth defects, and preterm birth, their causes are largely unknown. We seek to advance the use of social media for observational studi…

Epidemiology

Towards Detecting Inconsistencies in End-to-end Generated TODs

2026-07-10 · Tiziano Labruna, Giovanni Bonetta, Bernardo Magnini arxiv

Generative AI is profoundly transforming the core technologies behind conversational systems, shifting from component-based to end-to-end approaches. However, Large Language Models (LLMs) may still generate inconsistenci…

UAVs and Neural Networks for search and rescue missions

2023-10-09 · Hartmut Surmann, Artur Leinweber, Gerhard Senkowski, Julien Meine 외

In this paper, we present a method for detecting objects of interest, including cars, humans, and fire, in aerial images captured by unmanned aerial vehicles (UAVs) usually during vegetation fires. To achieve this, we us…

Data Augmentationobject-detectionObject Detection

Looking for COVID-19 misinformation in multilingual social media texts

2021-05-03 · Raj Ratn Pranesh, Mehrdad Farokhnejad, Ambesh Shekhar, Genoveva Vargas-Solar

This paper presents the Multilingual COVID-19 Analysis Method (CMTA) for detecting and observing the spread of misinformation about this disease within texts. CMTA proposes a data science (DS) pipeline that applies machi…

Misinformation