paper-with-me

홈 › Papers

OpenGloss: A Synthetic Encyclopedic Dictionary and Semantic Knowledge Graph

2025-11-23 · Michael J. Bommarito arxiv

We present OpenGloss, a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. OpenGloss contains 537K senses across 150K lexemes, on par with WordNet 3.1 and Open English WordNet, while providing more than four times as many sense definitions. These lexemes include 9.1M semantic edges, 1M usage examples, 3M collocations, and 60M words of encyclopedic content. Generated through a multi-agent procedural generation pipeline with schema-validated LLM outputs and automated quality assurance, the entire resource was produced in under one week for under $1,000. This demonstrates that structured generation can create comprehensive lexical resources at cost and time scales impractical for manual curation, enabling rapid iteration as foundation models improve. The resource addresses gaps in pedagogical applications by providing integrated content -- definitions, examples, collocations, encyclopedias, etymology -- that supports both vocabulary learning and natural language processing tasks. As a synthetically generated resource, OpenGloss reflects both the capabilities and limitations of current foundation models. The dataset is publicly available on Hugging Face under CC-BY 4.0, enabling researchers and educators to build upon and adapt this resource.

📄 PDF Abstract BibTeX arXiv:2511.18622

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Using BabelNet to Improve OOV Coverage in SMT

2016-05-01 · LREC 2016 5 · Jinhua Du, Andy Way, Andrzej Zydron

Out-of-vocabulary words (OOVs) are a ubiquitous and difficult problem in statistical machine translation (SMT). This paper studies different strategies of using BabelNet to alleviate the negative impact brought about by …

Domain AdaptationMachine TranslationTranslation

Multilinguality at Your Fingertips : BabelNet, Babelfy and Beyond !

2015-06-01 · JEPTALNRECITAL 2015 6 · Roberto Navigli

Multilinguality is a key feature of today{'}s Web, and it is this feature that we leverage and exploit in our research work at the Sapienza University of Rome{'}s Linguistic Computing Laboratory, which I am going to over…

Entity LinkingSemantic SimilaritySemantic Textual SimilarityWord Sense Disambiguation

Towards Building a Multilingual Sememe Knowledge Base: Predicting Sememes for BabelNet Synsets

2019-12-04 · Fanchao Qi, Liang Chang, Maosong Sun, Sicong Ouyang 외

A sememe is defined as the minimum semantic unit of human languages. Sememe knowledge bases (KBs), which contain words annotated with sememes, have been successfully applied to many NLP tasks. However, existing sememe KB…

Representing Multilingual Data as Linked Data: the Case of BabelNet 2.0

2014-05-01 · LREC 2014 5 · Maud Ehrmann, Francesco Cecconi, Daniele Vannella, John Philip McCrae 외

Recent years have witnessed a surge in the amount of semantic information published on the Web. Indeed, the Web of Data, a subset of the Semantic Web, has been increasing steadily in both volume and variety, transforming…

Enriching Ontologies with Encyclopedic Background Knowledge for Document Indexing

2016-03-21 · Posch Lisa

The rapidly increasing number of scientific documents available publicly on the Internet creates the challenge of efficiently organizing and indexing these documents. Due to the time consuming and tedious nature of manua…

BIG-bench Machine Learning