paper-with-me

홈 › Papers

SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions

2026-07-30 · Jennifer D'Souza, Sameer Sadruddin, Anisa Rula, Ana Bossler, Andrés Fullana, Enric Bas, Syed Ather, Defne Circi, Anlan Chen, L. Catherine Brinson, Alyssa Columbus, George Demetriou, Dongjun Jeong, Tarun Kumar, Frank Krüger, Sascha Genehr, Kai Budde-Sagert, Anamaria Leonescu, Francesco Lodola, Chiara Florindi, Gagana Balasubramanya Murthy, Samson Oluwapelumi Olagbile, Nazia Riasat, Yan Sha, Kevin Shen, Shaokai Yang arxiv

Scientific processes are often described in heterogeneous article discourse, with details needed for comparison, reproducibility, reuse, and automation dispersed across prose, tables, figures, protocols, and supplementary files. We present the first release of SciSchema.org, a multidisciplinary collection of 16 expert-annotated schemas spanning Biology & Biotechnology, Materials & Chemistry, Imaging & Measurement, Physics, and Psychology. Each schema defines reusable fields for describing process instances, including inputs, outputs, materials, instruments or software, parameters, conditions, procedural steps, measurements, and provenance-related information. The schemas were created through a human-in-the-loop schema-mining workflow in which large language models generated candidate structures from process specifications, scientific articles, and expert feedback, followed by domain-expert construction of final master schemas. The dataset contains final schemas in JSON Schema and SHACL formats, intermediate model-generated schemas, expert-feedback records, source-paper metadata, community-development materials, and analysis scripts. Technical validation assessed schema structure, development provenance, expert review, and syntactic conformance. The collection supports structured annotation, metadata enrichment, scientific knowledge graphs, information extraction, semantic publishing, and cross-study comparison.

📄 PDF Abstract BibTeX arXiv:2607.27955

Code (0)

등록된 구현이 없습니다.

Tasks

Information ExtractionKnowledge Graphs

Similar Papers 제목 키워드 기반

CookingSense: A Culinary Knowledgebase with Multidisciplinary Assertions

2024-05-01 · Donghee Choi, Mogan Gim, Donghyeon Park, Mujeen Sung 외

This paper introduces CookingSense, a descriptive collection of knowledge assertions in the culinary domain extracted from various sources, including web data, scientific papers, and recipes, from which knowledge coverin…

DescriptiveLanguage ModelingLanguage ModellingRetrieval

The STEM-ECR Dataset: Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources

2020-03-02 · LREC 2020 5 · Jennifer D'Souza, Anett Hoppe, Arthur Brack, Mohamad Yaser Jaradeh 외

We introduce the STEM (Science, Technology, Engineering, and Medicine) Dataset for Scientific Entity Extraction, Classification, and Resolution, version 1.0 (STEM-ECR v1.0). The STEM-ECR v1.0 dataset has been developed t…

Entity Extraction using GANEntity LinkingEntity ResolutionGeneral Classification+1

A Google-Proof Collection of French Winograd Schemas

2017-04-01 · WS 2017 4 · Pascal Amsili, Olga Seminck

This article presents the first collection of French Winograd Schemas. Winograd Schemas form anaphora resolution problems that can only be resolved with extensive world knowledge. For this reason the Winograd Schema Chal…

Coreference ResolutionWorld Knowledge

Overview of STEM Science as Process, Method, Material, and Data Named Entities

2022-05-24 · Jennifer D'Souza

We are faced with an unprecedented production in scholarly publications worldwide. Stakeholders in the digital libraries posit that the document-based publishing paradigm has reached the limits of adequacy. Instead, stru…

ArticlesAstronomyKnowledge GraphsNER

LLMs4SchemaDiscovery: A Human-in-the-Loop Workflow for Scientific Schema Mining with Large Language Models

2025-04-01 · Sameer Sadruddin, Jennifer D'Souza, Eleni Poupaki, Alex Watkins 외

Extracting structured information from unstructured text is crucial for modeling real-world processes, but traditional schema mining relies on semi-structured data, limiting scalability. This paper introduces schema-mine…