paper-with-me

홈 › Papers

A tale of two toolkits, report the first: benchmarking time series classification algorithms for correctness and efficiency

2019-09-12 · Anthony Bagnall, Franz Király, Markus Löning, Matthew Middlehurst, George Oastler

sktime is an open source, Python based, sklearn compatible toolkit for time series analysis developed by researchers at the University of East Anglia (UEA), University College London and the Alan Turing Institute. A key initial goal for sktime was to provide time series classification functionality equivalent to that available in a related java package, tsml, also developed at UEA. We describe the implementation of six such classifiers in sktime and compare them to their tsml equivalents. We demonstrate correctness through equivalence of accuracy on a range of standard test problems and compare the build time of the different implementations. We find that there is significant difference in accuracy on only one of the six algorithms we look at (Proximity Forest). This difference is causing us some pain in debugging. We found a much wider range of difference in efficiency. Again, this was not unexpected, but it does highlight ways both toolkits could be improved.

📄 PDF Abstract BibTeX arXiv:1909.05738

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingGeneral ClassificationTime SeriesTime Series AnalysisTime Series Classification

Similar Papers 제목 키워드 기반

A tale of two toolkits, report the third: on the usage and performance of HIVE-COTE v1.0

2020-04-13 · Anthony Bagnall, Michael Flynn, James Large, Jason Lines 외

The Hierarchical Vote Collective of Transformation-based Ensembles (HIVE-COTE) is a heterogeneous meta ensemble for time series classification. Since it was first proposed in 2016, the algorithm has undergone some minor …

Time SeriesTime Series AnalysisTime Series Classification

A tale of two toolkits, report the second: bake off redux. Chapter 1. dictionary based classifiers

2019-11-27 · Anthony Bagnall, James Large, Matthew Middlehurst

Time series classification (TSC) is the problem of learning labels from time dependent data. One class of algorithms is derived from a bag of words approach. A window is run along a series, the subseries is shortened and…

ClassificationGeneral ClassificationTime SeriesTime Series Analysis+1

MetaLead: A Comprehensive Human-Curated Leaderboard Dataset for Transparent Reporting of Machine Learning Experiments

2026-01-30 · Roelien C. Timmer, Necva Bölücü, Stephen Wan arxiv

Leaderboards are crucial in the machine learning (ML) domain for benchmarking and tracking progress. However, creating leaderboards traditionally demands significant manual effort. In recent years, efforts have been made…

Talent Hoarding in Organizations

2022-06-30 · Ingrid Haegele

Most organizations rely on managers to identify talented workers. However, managers who are evaluated on team performance have an incentive to hoard workers. This study provides the first empirical evidence of talent hoa…

State-of-the-Art Vietnamese Word Segmentation

2019-06-18 · Song Nguyen Duc Cong, Quoc Hung Ngo, Rachsuda Jiamthapthaksin

Word segmentation is the first step of any tasks in Vietnamese language processing. This paper reviews stateof-the-art approaches and systems for word segmentation in Vietnamese. To have an overview of all stages from bu…

BIG-bench Machine LearningSegmentationVietnamese Word Segmentation