MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization
Recently, large-scale datasets have vastly facilitated the development in nearly all domains of Natural Language Processing. However, there is currently no cross-task dataset in NLP, which hinders the development of multi-task learning. We propose MATINF, the first jointly labeled large-scale dataset for classification, question answering and summarization. MATINF contains 1.07 million question-answer pairs with human-labeled categories and user-generated question descriptions. Based on such rich information, MATINF is applicable for three major NLP tasks, including classification, question answering, and summarization. We benchmark existing methods and a novel multi-task baseline over MATINF to inspire further research. Our comprehensive comparison and experiments over MATINF and other datasets demonstrate the merits held by MATINF.
Code (1)
Tasks
ClassificationGeneral ClassificationMulti-Task LearningQuestion AnsweringSimilar Papers 제목 키워드 기반
Materials Informatics Transformer: A Language Model for Interpretable Materials Properties Prediction
Recently, the remarkable capabilities of large language models (LLMs) have been illustrated across a variety of research domains such as natural language processing, computer vision, and molecular modeling. We extend thi…
Language ModelingLanguage ModellingPredictionProperty PredictionLarge-Scale Land Cover Mapping with Fine-Grained Classes via Class-Aware Semi-Supervised Semantic Segmentation
Semi-supervised learning has attracted increasing attention in the large-scale land cover mapping task. However, existing methods overlook the potential to alleviate the class imbalance problem by selecting a suitabl…
Semantic SegmentationSemi-Supervised Semantic SegmentationBertGCN: Transductive Text Classification by Combining GCN and BERT
In this work, we propose BertGCN, a model that combines large scale pretraining and transductive learning for text classification. BertGCN constructs a heterogeneous graph over the dataset and represents documents as nod…
Classificationtext-classificationText ClassificationTransductive LearningToward Deep Supervised Anomaly Detection: Reinforcement Learning from Partially Labeled Anomaly Data
We consider the problem of anomaly detection with a small set of partially labeled anomaly examples and a large-scale unlabeled dataset. This is a common scenario in many important applications. Existing related methods …
Anomaly DetectionDeep Reinforcement Learningreinforcement-learningReinforcement Learning (RL)+1Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained Computation
We propose and study a realistic Continual Learning (CL) setting where learning algorithms are granted a restricted computational budget per time step while training. We apply this setting to large-scale semi-supervised …
Continual Learning