SPLIT: Smart Preprocessing (Quasi) Language Independent Tool
Text preprocessing is an important and necessary task for all NLP applications. A simple variation in any preprocessing step may drastically affect the final results. Moreover replicability and comparability, as much as feasible, is one of the goals of our scientific enterprise, thus building systems that can ensure the consistency in our various pipelines would contribute significantly to our goals. The problem has become quite pronounced with the abundance of NLP tools becoming more and more available yet with different levels of specifications. In this paper, we present a dynamic unified preprocessing framework and tool, SPLIT, that is highly configurable based on user requirements which serves as a preprocessing tool for several tools at once. SPLIT aims to standardize the implementations of the most important preprocessing steps by allowing for a unified API that could be exchanged across different researchers to ensure complete transparency in replication. The user is able to select the required preprocessing tasks among a long list of preprocessing steps. The user is also able to specify the order of execution which in turn affects the final preprocessing output.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A Controlled Visual-Backbone Benchmark for Multimodal Short-Term Solar Irradiance Forecasting
Sky-image irradiance studies often compare forecasting systems in which the image encoder, temporal model, fusion block, target definition, and training recipe all change together. We use a narrower protocol: the multimo…
Solar Irradiance ForecastingResolution-independent meshes of super pixels
The over-segmentation into superpixels is an important preprocessing step to smartly compress the input size and speed up higher level tasks. A superpixel was traditionally considered as a small cluster of square-based p…
ClusteringSegmentationSuperpixelsSmartSplit: Latency-Energy-Memory Optimisation for CNN Splitting on Smartphone Environment
Artificial Intelligence has now taken centre stage in the smartphone industry owing to the need of bringing all processing close to the user and addressing privacy concerns. Convolution Neural Networks (CNNs), which are …
Design Choices in Splitting-Based Self-Supervised Sparse-View CT Reconstruction
Self-supervised data splitting has emerged as a promising paradigm for sparse-view CT reconstruction, enabling training from incomplete measurements without fully sampled ground truth. However, the influence of key desig…
Smart Data driven Decision Trees Ensemble Methodology for Imbalanced Big Data
Differences in data size per class, also known as imbalanced data distribution, have become a common problem affecting data quality. Big Data scenarios pose a new challenge to traditional imbalanced classification algori…
BIG-bench Machine LearningGeneral Classificationimbalanced classification