paper-with-me

Papers

Sublanguage Corpus Analysis Toolkit: A tool for assessing the representativeness and sublanguage characteristics of corpora

2014-05-01 · LREC 2014 5 · Irina Temnikova, William A. Baumgartner Jr., Negacy D. Hailu, Ivelina Nikolova, Tony McEnery, Adam Kilgarriff, Galia Angelova, K. Bretonnel Cohen

Sublanguages are varieties of language that form “subsets” of the general language, typically exhibiting particular types of lexical, semantic, and other restrictions and deviance. SubCAT, the Sublanguage Corpus Analysis Toolkit, assesses the representativeness and closure properties of corpora to analyze the extent to which they are either sublanguages, or representative samples of the general language. The current version of SubCAT contains scripts and applications for assessing lexical closure, morphological closure, sentence type closure, over-represented words, and syntactic deviance. Its operation is illustrated with three case studies concerning scientific journal articles, patents, and clinical records. Materials from two language families are analyzed―English (Germanic), and Bulgarian (Slavic). The software is available at sublanguage.sourceforge.net under a liberal Open Source license.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesSentence

Similar Papers 제목 키워드 기반

SuperCAT: The (New and Improved) Corpus Analysis Toolkit

2016-05-01 · LREC 2016 5 · K. Bretonnel Cohen, William A. Baumgartner Jr., Irina Temnikova

This paper reports SuperCAT, a corpus analysis toolkit. It is a radical extension of SubCAT, the Sublanguage Corpus Analysis Toolkit, from sublanguage analysis to corpus analysis in general. The idea behind SuperCAT is t…

Global Pre-ordering for Improving Sublanguage Translation

2016-12-01 · WS 2016 12 · Masaru Fuji, Masao Utiyama, Eiichiro Sumita, Yuji Matsumoto

When translating formal documents, capturing the sentence structure specific to the sublanguage is extremely necessary to obtain high-quality translations. This paper proposes a novel global reordering method with partic…

Machine TranslationSentenceTranslation

Open corpora and toolkit for assessing text readability in French

2022-06-01 · READI (LREC) 2022 6 · Nicolas Hernandez, Nabil Oulbaz, Tristan Faine

Measuring the linguistic complexity or assessing the readability of spoken or written productions has been the concern of several researchers in pedagogy and (foreign) language teaching for decades. Researchers study for…

Text Simplification

GeBioToolkit: Automatic Extraction of Gender-Balanced Multilingual Corpus of Wikipedia Biographies

2019-12-10 · LREC 2020 5 · Marta R. Costa-jussà, Pau Li Lin, Cristina España-Bonet

We introduce GeBioToolkit, a tool for extracting multilingual parallel corpora at sentence level, with document and gender information from Wikipedia biographies. Despite thegender inequalitiespresent in Wikipedia, the t…

Sentence

Pimlico: A toolkit for corpus-processing pipelines and reproducible experiments

2020-11-01 · EMNLP (NLPOSS) 2020 11 · Mark Granroth-Wilding

We present Pimlico, an open source toolkit for building pipelines for processing large corpora. It is especially focused on processing linguistic corpora and provides wrappers around existing, widely used NLP tools. A pa…