paper-with-me

Papers

Global Intelligent Content: Active Curation of Language Resources using Linked Data

2014-05-01 · LREC 2014 5 · David Lewis, Rob Brennan, Leroy Finn, Dominic Jones, Alan Meehan, Declan O{'}Sullivan, Sebastian Hellmann, Felix Sasaki

As language resources start to become available in linked data formats, it becomes relevant to consider how linked data interoperability can play a role in active language processing workflows as well as for more static language resource publishing. This paper proposes that linked data may have a valuable role to play in tracking the use and generation of language resources in such workflows in order to assess and improve the performance of the language technologies that use the resources, based on feedback from the human involvement typically required within such processes. We refer to this as Active Curation of the language resources, since it is performed systematically over language processing workflows to continuously improve the quality of the resource in specific applications, rather than via dedicated curation steps. We use modern localisation workflows, i.e. assisted by machine translation and text analytics services, to explain how linked data can support such active curation. By referencing how a suitable linked data vocabulary can be assembled by combining existing linked data vocabularies and meta-data from other multilingual content processing annotations and tool exchange standards we aim to demonstrate the relative ease with which active curation can be deployed more broadly.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalMachine TranslationTranslation

Similar Papers 제목 키워드 기반

QURATOR: Innovative Technologies for Content and Data Curation

2020-04-25 · Georg Rehm, Peter Bourgonje, Stefanie Hegele, Florian Kintzel 외

In all domains and sectors, the demand for intelligent systems to support the processing and generation of digital content is rapidly increasing. The availability of vast amounts of content and the pressure to publish ne…

ShennongAlpha: an AI-driven sharing and collaboration platform for intelligent curation, acquisition, and translation of natural medicinal material knowledge

2023-12-27 · Zijie Yang, Yongjing Yin, Chaojun Kong, Tiange Chi 외

Natural Medicinal Materials (NMMs) have a long history of global clinical applications and a wealth of records and knowledge. Although NMMs are a major source for drug discovery and clinical application, the utilization …

Drug DiscoveryMachine TranslationManagementTranslation

Oasis: Data Curation and Assessment System for Pretraining of Large Language Models

2023-11-21 · Tong Zhou, Yubo Chen, Pengfei Cao, Kang Liu 외

Data is one of the most critical elements in building a large language model. However, existing systems either fail to customize a corpus curation pipeline or neglect to leverage comprehensive corpus assessment for itera…

Language ModelingLanguage ModellingLarge Language Model

Finder: A Multimodal AI-Powered Search Framework for Pharmaceutical Data Retrieval

2026-01-06 · Suyash Mishra, Srikanth Patil, Satyanarayan Pati, Sagar Sahu 외 arxiv

AI is transforming pharmaceutical search, where traditional systems struggle with multimodal content and manual curation. Finder is a scalable AI-powered framework that unifies retrieval across text, images, audio, and v…

ATHENA: Accelerated Multi-Task Heterogeneous Influence Functions for Robot Data Curation

2026-06-15 · Tao Xu, Jiaxin Wang, Runhao Zhang, Jiayi Guan 외 arxiv

In robot imitation learning, influence functions provide a principled approach to quantify each demonstration's effect on robot task outcomes, yet scaling them to billion-parameter Vision-Language-Action (VLA) models is …