paper-with-me

Papers

Extrinsic Corpus Evaluation with a Collocation Dictionary Task

2014-05-01 · LREC 2014 5 · Adam Kilgarriff, Pavel Rychl{\'y}, Milo{\v{s}} Jakub{\'\i}{\v{c}}ek, Vojt{\v{e}}ch Kov{\'a}{\v{r}}, V{\'\i}t Baisa, Lucia Kocincov{\'a}

The NLP researcher or application-builder often wonders {`}what corpus should I use, or should I build one of my own? If I build one of my own, how will I know if I have done a good job?{''} Currently there is very little help available for them. They are in need of a framework for evaluating corpora. We develop such a framework, in relation to corpora which aim for good coverage of general language{'}. The task we set is automatic creation of a publication-quality collocations dictionary. For a sample of 100 headwords of Czech and 100 of English, we identify a gold standard dataset of (ideally) all the collocations that should appear for these headwords in such a dictionary. The datasets are being made available alongside this paper. We then use them to determine precision and recall for a range of corpora, with a range of parameters.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Acquiring Social Knowledge about Personality and Driving-related Behavior

2020-05-01 · LREC 2020 5 · Ritsuko Iwai, Daisuke Kawahara, Takatsune Kumada, Sadao Kurohashi

In this paper, we introduce our psychological approach to collect human-specific social knowledge from a text corpus, using NLP techniques. It is often not explicitly described but shared among people, which we call soci…

Extending WordNet with Fine-Grained Collocational Information via Supervised Distributional Learning

2016-12-01 · COLING 2016 12 · Luis Espinosa-Anke, Jose Camacho-Collados, Sara Rodr{\'\i}guez-Fern{\'a}ndez, Horacio Saggion 외

WordNet is probably the best known lexical resource in Natural Language Processing. While it is widely regarded as a high quality repository of concepts and semantic relations, updating and extending it manually is costl…

Machine TranslationSemantic Textual SimilaritySentiment AnalysisText Generation+3

Urban Dictionary Embeddings for Slang NLP Applications

2020-05-01 · LREC 2020 5 · Steven Wilson, Walid Magdy, Barbara McGillivray, Kiran Garimella 외

The choice of the corpus on which word embeddings are trained can have a sizable effect on the learned representations, the types of analyses that can be performed with them, and their utility as features for machine lea…

ClusteringSarcasm DetectionSemantic SimilaritySemantic Textual Similarity+2

Gigafida 2.0: The Reference Corpus of Written Standard Slovene

2020-05-01 · LREC 2020 5 · Simon Krek, {\v{S}}pela Arhar Holdt, Toma{\v{z}} Erjavec, Jaka {\v{C}}ibej 외

We describe a new version of the Gigafida reference corpus of Slovene. In addition to updating the corpus with new material and annotating it with better tools, the focus of the upgrade was also on its transformation fro…

Automatic detection of unexpected/erroneous collocations in learner corpus

2020-12-01 · COLING (MWE) 2020 12 · Jen-Yu Li, Thomas Gaillat

This research investigates the collocational errors made by English learners in a learner corpus. It focuses on the extraction of unexpected collocations. A system was proposed and implemented with open source toolkit. F…