paper-with-me

홈 › Papers

A Comparative Evaluation of Curriculum Learning with Filtering and Boosting

2013-12-17 · Michael R. Smith, Tony Martinez

Not all instances in a data set are equally beneficial for inferring a model of the data. Some instances (such as outliers) are detrimental to inferring a model of the data. Several machine learning techniques treat instances in a data set differently during training such as curriculum learning, filtering, and boosting. However, an automated method for determining how beneficial an instance is for inferring a model of the data does not exist. In this paper, we present an automated method that orders the instances in a data set by complexity based on the their likelihood of being misclassified (instance hardness). The underlying assumption of this method is that instances with a high likelihood of being misclassified represent more complex concepts in a data set. Ordering the instances in a data set allows a learning algorithm to focus on the most beneficial instances and ignore the detrimental ones. We compare ordering the instances in a data set in curriculum learning, filtering and boosting. We find that ordering the instances significantly increases classification accuracy and that filtering has the largest impact on classification accuracy. On a set of 52 data sets, ordering the instances increases the average accuracy from 81% to 84%.

📄 PDF Abstract BibTeX arXiv:1312.4986

Code (0)

등록된 구현이 없습니다.

Tasks

General Classification

Similar Papers 제목 키워드 기반

Efficient Pre-training of Masked Language Model via Concept-based Curriculum Masking

2022-12-15 · Mingyu Lee, Jun-Hyung Park, Junho Kim, Kang-Min Kim 외

Masked language modeling (MLM) has been widely used for pre-training effective bidirectional representations, but incurs substantial training costs. In this paper, we propose a novel concept-based curriculum masking (CCM…

Language ModelingLanguage ModellingMasked Language Modeling

Categorizing Comparative Sentences

2018-09-17 · WS 2019 8 · Alexander Panchenko, Alexander Bondarenko, Mirco Franzek, Matthias Hagen 외

We tackle the tasks of automatically identifying comparative sentences and categorizing the intended preference (e.g., "Python has better NLP libraries than MATLAB" => (Python, better, MATLAB). To this end, we manually a…

Argument MiningSentenceSentence Embeddings

Reinforcement Learning based Curriculum Optimization for Neural Machine Translation

2019-02-28 · NAACL 2019 6 · Gaurav Kumar, George Foster, Colin Cherry, Maxim Krikun

We consider the problem of making efficient use of heterogeneous training data in neural machine translation (NMT). Specifically, given a training dataset with a sentence-level feature such as noise, we seek an optimal c…

Machine TranslationNMTreinforcement-learningReinforcement Learning+3

PCC: Paraphrasing with Bottom-k Sampling and Cyclic Learning for Curriculum Data Augmentation

2022-08-17 · Hongyuan Lu, Wai Lam

Curriculum Data Augmentation (CDA) improves neural models by presenting synthetic data with increasing difficulties from easy to hard. However, traditional CDA simply treats the ratio of word perturbation as the difficul…

Data AugmentationDialogue GenerationFew-Shot Text ClassificationParaphrase Generation+2

ALAS: Autonomous Learning Agent for Self-Updating Language Models

2025-08-14 · Dhruv Atreja arxiv

Large language models (LLMs) often have a fixed knowledge cutoff, limiting their accuracy on emerging information. We present ALAS (Autonomous Learning Agent System), a modular pipeline that continuously updates an LLM's…

Continual LearningQuestion Answering