Tunable Distortion Limits and Corpus Cleaning for SMT
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModellingMachine TranslationSimilar Papers 제목 키워드 기반
Developing an efficient corpus using Ensemble Data cleaning approach
Despite the observable benefit of Natural Language Processing (NLP) in processing a large amount of textual medical data within a limited time for information retrieval, a handful of research efforts have been devoted to…
Information RetrievalData Cleaning Tools for Token Classification Tasks
Human-in-the-loop systems for cleaning NLP training data rely on automated sieves to isolate potentially-incorrect labels for manual review. We have developed a novel technique for flagging potentially-incorrect labels w…
Classificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+3Knowledge Enhanced Sports Game Summarization
Sports game summarization aims at generating sports news from live commentaries. However, existing datasets are all constructed through automated collection and cleaning processes, resulting in a lot of noise. Besides, c…
tsrobprep - an R package for robust preprocessing of time series data
Data cleaning is a crucial part of every data analysis exercise. Yet, the currently available R packages do not provide fast and robust methods for cleaning and preparation of time series data. The open source package ts…
ClusteringImputationMissing ValuesOutlier Detection+2Building a 70 billion word corpus of English from ClueWeb
This work describes the process of creation of a 70 billion word text corpus of English. We used an existing language resource, namely the ClueWeb09 dataset, as source for the corpus data. Processing such a vast amount o…
Machine TranslationManagementPart-Of-Speech TaggingSpeech Recognition