Building a reference lexicon for countability in English
The present paper describes the construction of a resource to determine the lexical preference class of a large number of English noun-senses ({\$}{\textbackslash}approx{\$} 14,000) with respect to the distinction between mass and count interpretations. In constructing the lexicon, we have employed a questionnaire-based approach based on existing resources such as the Open ANC ({\textbackslash}url{http://www.anc.org}) and WordNet {\textbackslash}cite{Miller95}. The questionnaire requires annotators to answer six questions about a noun-sense pair. Depending on the answers, a given noun-sense pair can be assigned to fine-grained noun classes, spanning the area between count and mass. The reference lexicon contains almost 14,000 noun-sense pairs. An initial data set of 1,000 has been annotated together by four native speakers, while the remaining 12,800 noun-sense pairs have been annotated in parallel by two annotators each. We can confirm the general feasibility of the approach by reporting satisfactory values between 0.694 and 0.755 in inter-annotator agreement using Krippendorff{'}s {\$}{\textbackslash}alpha{\$}.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A sense-based lexicon of count and mass expressions: The Bochum English Countability Lexicon
The present paper describes the current release of the Bochum English Countability Lexicon (BECL 2.1), a large empirical database consisting of lemmata from Open ANC (http://www.anc.org) with added senses from WordNet (F…
Semantic SimilaritySemantic Textual SimilarityNon-native English lexicon creation for bilingual speech synthesis
Bilingual English speakers speak English as one of their languages. Their English is of a non-native kind, and their conversations are of a code-mixed fashion. The intelligibility of a bilingual text-to-speech (TTS) syst…
Speech Synthesistext-to-speechText to SpeechTowards a Computational Lexicon for Moroccan Darija: Words, Idioms, and Constructions
In this paper, we explore the challenges of building a computational lexicon for Moroccan Darija (MD), an Arabic dialect spoken by over 32 million people worldwide but which only recently has begun appearing frequently i…
Machine TranslationHABLex: Human Annotated Bilingual Lexicons for Experiments in Machine Translation
Bilingual lexicons are valuable resources used by professional human translators. While these resources can be easily incorporated in statistical machine translation, it is unclear how to best do so in the neural framewo…
Machine TranslationTranslationSynSemClass Linked Lexicon: Mapping Synonymy between Languages
This paper reports on an extended version of a synonym verb class lexicon, newly called SynSemClass (formerly CzEngClass). This lexicon stores cross-lingual semantically similar verb senses in synonym classes extracted f…