Data Augmentation Techniques for Process Extraction from Scientific Publications
We present data augmentation techniques for process extraction tasks in scientific publications. We cast the process extraction task as a sequence labeling task where we identify all the entities in a sentence and label them according to their process-specific roles. The proposed method attempts to create meaningful augmented sentences by utilizing (1) process-specific information from the original sentence, (2) role label similarity, and (3) sentence similarity. We demonstrate that the proposed methods substantially improve the performance of the process extraction model trained on chemistry domain datasets, up to 12.3 points improvement in performance accuracy (F-score). The proposed methods could potentially reduce overfitting as well, especially when training on small datasets or in a low-resource setting such as in chemistry and other scientific domains.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationSentenceSentence SimilaritySimilar Papers 제목 키워드 기반
Leveraging Data Augmentation for Process Information Extraction
Business Process Modeling projects often require formal process models as a central component. High costs associated with the creation of such formal process models motivated many different fields of research aimed at au…
Data AugmentationRelation ExtractionPGA-SciRE: Harnessing LLM on Data Augmentation for Enhancing Scientific Relation Extraction
Relation Extraction (RE) aims at recognizing the relation between pairs of entities mentioned in a text. Advances in LLMs have had a tremendous impact on NLP. In this work, we propose a textual data augmentation framewor…
Data AugmentationRelationRelation ExtractionSentenceExploring Data Augmentation and Resampling Strategies for Transformer-Based Models to Address Class Imbalance in AI Scoring of Scientific Explanations in NGSS Classroom
Automated scoring of students' scientific explanations offers the potential for immediate, accurate feedback, yet class imbalance in rubric categories particularly those capturing advanced reasoning remains a challenge. …
Text ClassificationData AugmentationMaking Invisible Visible: Data-Driven Seismic Inversion with Spatio-temporally Constrained Data Augmentation
Deep learning and data-driven approaches have shown great potential in scientific domains. The promise of data-driven techniques relies on the availability of a large volume of high-quality training datasets. Due to the …
Data AugmentationSeismic ImagingSeismic InversionOhioState at SemEval-2018 Task 7: Exploiting Data Augmentation for Relation Classification in Scientific Papers using Piecewise Convolutional Neural Networks
We describe our system for SemEval-2018 Shared Task on Semantic Relation Extraction and Classification in Scientific Papers where we focus on the Classification task. Our simple piecewise convolution neural encoder perfo…
ClassificationData AugmentationGeneral ClassificationRelation Classification+1