paper-with-me

Papers

WildlifeDatasets: An open-source toolkit for animal re-identification

2023-11-15 · Vojtěch Čermák, Lukas Picek, Lukáš Adam, Kostas Papafitsoros

In this paper, we present WildlifeDatasets (https://github.com/WildlifeDatasets/wildlife-datasets) - an open-source toolkit intended primarily for ecologists and computer-vision / machine-learning researchers. The WildlifeDatasets is written in Python, allows straightforward access to publicly available wildlife datasets, and provides a wide variety of methods for dataset pre-processing, performance analysis, and model fine-tuning. We showcase the toolkit in various scenarios and baseline experiments, including, to the best of our knowledge, the most comprehensive experimental comparison of datasets and methods for wildlife re-identification, including both local descriptors and deep learning approaches. Furthermore, we provide the first-ever foundation model for individual re-identification within a wide range of species - MegaDescriptor - that provides state-of-the-art performance on animal re-identification datasets and outperforms other pre-trained models such as CLIP and DINOv2 by a significant margin. To make the model available to the general public and to allow easy integration with any existing wildlife monitoring applications, we provide multiple MegaDescriptor flavors (i.e., Small, Medium, and Large) through the HuggingFace hub (https://huggingface.co/BVRA).

📄 PDF Abstract BibTeX arXiv:2311.09118

Code (2)

wildlifedatasets/wildlife-datasets 공식 구현
wildlifedatasets/wildlife-tools 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

SeaTurtleID2022: A long-span dataset for reliable sea turtle re-identification

2023-11-09 · Lukáš Adam, Vojtěch Čermák, Kostas Papafitsoros, Lukáš Picek

This paper introduces the first public large-scale, long-span dataset with sea turtle photographs captured in the wild -- SeaTurtleID2022 (https://www.kaggle.com/datasets/wildlifedatasets/seaturtleid2022). The dataset co…

BenchmarkingInstance SegmentationSegmentationSemantic Segmentation

SeaTurtleID2022: A long-span dataset for reliable sea turtle re-identification

2022-11-18 · Lukáš Adam, Vojtěch Čermák, Kostas Papafitsoros, Lukáš Picek

This paper introduces the first public large-scale, long-span dataset with sea turtle photographs captured in the wild -- \href{https://www.kaggle.com/datasets/wildlifedatasets/seaturtleid2022}{SeaTurtleID2022}. The data…

BenchmarkingInstance SegmentationSegmentationSemantic Segmentation

MAPA Project: Ready-to-Go Open-Source Datasets and Deep Learning Technology to Remove Identifying Information from Text Documents

2022-06-01 · LEGAL (LREC) 2022 6 · Victoria Arranz, Khalid Choukri, Montse Cuadros, Aitor García Pablos 외

This paper presents the outcomes of the MAPA project, a set of annotated corpora for 24 languages of the European Union and an open-source customisable toolkit able to detect and substitute sensitive information in text …

De-identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

The Multilingual Anonymisation Toolkit for Public Administrations (MAPA) Project

2020-11-01 · EAMT 2020 11 · Ēriks Ajausks, Victoria Arranz, Laurent Bié, Aleix Cerdà-i-Cucó 외

We describe the MAPA project, funded under the Connecting Europe Facility programme, whose goal is the development of an open-source de-identification toolkit for all official European Union languages. It will be develop…

De-identification

A Software Toolkit for Pre-processing Sign Language Video Streams

2022-06-01 · SLTAT (LREC) 2022 6 · Fabrizio Nunnari

We present the requirements, design guidelines, and the software architecture of an open-source toolkit dedicated to the pre-processing of sign language video material. The toolkit is a collection of functions and comman…