paper-with-me

홈 › Papers

NLTK: The Natural Language Toolkit

2002-05-17 · Edward Loper, Steven Bird

NLTK, the Natural Language Toolkit, is a suite of open source program modules, tutorials and problem sets, providing ready-to-use computational linguistics courseware. NLTK covers symbolic and statistical natural language processing, and is interfaced to annotated corpora. Students augment and replace existing components, learn structured programming by example, and manipulate sophisticated models from the outset.

📄 PDF Abstract BibTeX arXiv:cs/0205028

Code (1)

napakalas/NLIMED tf

Tasks

Multi-Label Text Classification

Similar Papers 제목 키워드 기반

EstNLTK - NLP Toolkit for Estonian

2016-05-01 · LREC 2016 5 · Siim Orasmaa, Timo Petmanson, Alex Tkachenko, er 외

Although there are many tools for natural language processing tasks in Estonian, these tools are very loosely interoperable, and it is not easy to build practical applications on top of them. In this paper, we introduce …

Morphological Analysisnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

iNLTK: Natural Language Toolkit for Indic Languages

2020-09-26 · EMNLP (NLPOSS) 2020 11 · Gaurav Arora

We present iNLTK, an open-source NLP library consisting of pre-trained language models and out-of-the-box support for Data Augmentation, Textual Similarity, Sentence Embeddings, Word Embeddings, Tokenization and Text Gen…

Data AugmentationParaphrase GenerationSentenceSentence Embeddings+4

Practical Approach on Implementation of WordNets for South African Languages

2021-01-01 · EACL (GWC) 2021 1 · Tshephisho Joseph Sefara, Tumisho Billson Mokgonyane, Vukosi Marivate

This paper proposes the implementation of WordNets for five South African languages, namely, Sepedi, Setswana, Tshivenda, isiZulu and isiXhosa to be added to open multilingual WordNets (OMW) on natural language toolkit (…

EstNLTK 1.6: Remastered Estonian NLP Pipeline

2020-05-01 · LREC 2020 5 · Sven Laur, Siim Orasmaa, Dage S{\"a}rg, Paul Tammo

The goal of the EstNLTK Python library is to provide a unified programming interface for natural language processing in Estonian. As such, previous versions of the library have been immensely successful both in academic …

Morphological Analysis

Text Normalization for Low-Resource Languages of Africa

2021-03-29 · Andrew Zupon, Evan Crew, Sandy Ritchie

Training data for machine learning models can come from many different sources, which can be of dubious quality. For resource-rich languages like English, there is a lot of data available, so we can afford to throw out t…

Text Normalization