paper-with-me

Papers

Harmonizing Metadata of Language Resources for Enhanced Querying and Accessibility

2025-01-09 · Zixuan Liang

This paper addresses the harmonization of metadata from diverse repositories of language resources (LRs). Leveraging linked data and RDF techniques, we integrate data from multiple sources into a unified model based on DCAT and META-SHARE OWL ontology. Our methodology supports text-based search, faceted browsing, and advanced SPARQL queries through Linghub, a newly developed portal. Real user queries from the Corpora Mailing List (CML) were evaluated to assess Linghub capability to satisfy actual user needs. Results indicate that while some limitations persist, many user requests can be successfully addressed. The study highlights significant metadata issues and advocates for adherence to open vocabularies and standards to enhance metadata harmonization. This initial research underscores the importance of API-based access to LRs, promoting machine usability and data subset extraction for specific purposes, paving the way for more efficient and standardized LR utilization.

📄 PDF Abstract BibTeX arXiv:2501.05606

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards automation in using multi-modal language resources: compatibility and interoperability for multi-modal features in Kachako

2012-05-01 · LREC 2012 5 · Yoshinobu Kano

Use of language resources including annotated corpora and tools is not easy for users, as it requires expert knowledge to determine which resources are compatible and interoperable. Sometimes it requires programming skil…

Use Case: Romanian Language Resources in the LOD Paradigm

2022-06-01 · LDL (ACL) 2022 6 · Verginica Barbu Mititelu, Elena Irimia, Vasile Pais, Andrei-Marius Avram 외

In this paper, we report on (i) the conversion of Romanian language resources to the Linked Open Data specifications and requirements, on (ii) their publication and (iii) interlinking with other language resources (for R…

Word Embeddings

FISHNET: Financial Intelligence from Sub-querying, Harmonizing, Neural-Conditioning, Expert Swarms, and Task Planning

2024-10-25 · Nicole Cho, Nishan Srishankar, Lucas Cecchi, William Watson

Financial intelligence generation from vast data sources has typically relied on traditional methods of knowledge-graph construction or database engineering. Recently, fine-tuned financial domain-specific Large Language …

graph constructionRAGTask Planning

Harmonizing Different Lemmatization Strategies for Building a Knowledge Base of Linguistic Resources for Latin

2019-08-01 · WS 2019 8 · Francesco Mambrini, Marco Passarotti

The interoperability between lemmatized corpora of Latin and other resources that use the lemma as indexing key is hampered by the multiple lemmatization strategies that different projects adopt. In this paper we discuss…

LEMMALemmatization

QQ: A Language Metadata Toolkit for Multilingual NLP

2026-02-28 · Wessel Poelman, Yiyi Chen, Miryam de Lhoneux arxiv

Multilingual NLP research increasingly involves hundreds or thousands of languages across different datasets. Managing, discovering, and reporting language metadata becomes a common hurdle at these scales. We present QQ,…