paper-with-me

홈 › Papers

Tunable Distortion Limits and Corpus Cleaning for SMT

2013-08-01 · WS 2013 8 · Sara Stymne, Christian Hardmeier, J{\"o}rg Tiedemann, Joakim Nivre
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModellingMachine Translation

Similar Papers 제목 키워드 기반

Developing an efficient corpus using Ensemble Data cleaning approach

2024-06-02 · Md Taimur Ahad

Despite the observable benefit of Natural Language Processing (NLP) in processing a large amount of textual medical data within a limited time for information retrieval, a handful of research efforts have been devoted to…

Information Retrieval

Data Cleaning Tools for Token Classification Tasks

2021-06-01 · NAACL (DaSH) 2021 6 · Karthik Muthuraman, Frederick Reiss, Hong Xu, Bryan Cutler 외

Human-in-the-loop systems for cleaning NLP training data rely on automated sieves to isolate potentially-incorrect labels for manual review. We have developed a novel technique for flagging potentially-incorrect labels w…

Classificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+3

Knowledge Enhanced Sports Game Summarization

2021-11-24 · Jiaan Wang, Zhixu Li, Tingyi Zhang, Duo Zheng 외

Sports game summarization aims at generating sports news from live commentaries. However, existing datasets are all constructed through automated collection and cleaning processes, resulting in a lot of noise. Besides, c…

tsrobprep - an R package for robust preprocessing of time series data

2021-04-26 · Michał Narajewski, Jens Kley-Holsteg, Florian Ziel

Data cleaning is a crucial part of every data analysis exercise. Yet, the currently available R packages do not provide fast and robust methods for cleaning and preparation of time series data. The open source package ts…

ClusteringImputationMissing ValuesOutlier Detection+2

Building a 70 billion word corpus of English from ClueWeb

2012-05-01 · LREC 2012 5 · Jan Pomik{\'a}lek, Milo{\v{s}} Jakub{\'\i}{\v{c}}ek, Pavel Rychl{\'y}

This work describes the process of creation of a 70 billion word text corpus of English. We used an existing language resource, namely the ClueWeb09 dataset, as source for the corpus data. Processing such a vast amount o…

Machine TranslationManagementPart-Of-Speech TaggingSpeech Recognition