paper-with-me

홈 › Papers

PyGraft: Configurable Generation of Synthetic Schemas and Knowledge Graphs at Your Fingertips

2023-09-07 · Nicolas Hubert, Pierre Monnin, Mathieu d'Aquin, Davy Monticolo, Armelle Brun

Knowledge graphs (KGs) have emerged as a prominent data representation and management paradigm. Being usually underpinned by a schema (e.g., an ontology), KGs capture not only factual information but also contextual knowledge. In some tasks, a few KGs established themselves as standard benchmarks. However, recent works outline that relying on a limited collection of datasets is not sufficient to assess the generalization capability of an approach. In some data-sensitive fields such as education or medicine, access to public datasets is even more limited. To remedy the aforementioned issues, we release PyGraft, a Python-based tool that generates highly customized, domain-agnostic schemas and KGs. The synthesized schemas encompass various RDFS and OWL constructs, while the synthesized KGs emulate the characteristics and scale of real-world KGs. Logical consistency of the generated resources is ultimately ensured by running a description logic (DL) reasoner. By providing a way of generating both a schema and KG in a single pipeline, PyGraft's aim is to empower the generation of a more diverse array of KGs for benchmarking novel approaches in areas such as graph-based machine learning (ML), or more generally KG processing. In graph-based ML in particular, this should foster a more holistic evaluation of model performance and generalization capability, thereby going beyond the limited collection of available benchmarks. PyGraft is available at: https://github.com/nicolas-hbt/pygraft.

📄 PDF Abstract BibTeX arXiv:2309.03685

Code (1)

nicolas-hbt/pygraft 공식 구현

Tasks

BenchmarkingKnowledge Graphs

Similar Papers 제목 키워드 기반

Schema Generation for Large Knowledge Graphs Using Large Language Models

2025-06-04 · Bohui Zhang, Yuan He, Lydia Pintscher, Albert Meroño Peñuela 외

Schemas are vital for ensuring data quality in the Semantic Web and natural language processing. Traditionally, their creation demands substantial involvement from knowledge engineers and domain experts. Leveraging the i…

Knowledge Graphs

Harvesting Event Schemas from Large Language Models

2023-05-12 · Jialong Tang, Hongyu Lin, Zhuoqun Li, Yaojie Lu 외

Event schema provides a conceptual, structural and formal language to represent events and model the world event knowledge. Unfortunately, it is challenging to automatically induce high-quality and high-coverage event sc…

Diversity

We are what we repeatedly do: Inducing and deploying habitual schemas in persona-based responses

2023-10-10 · Benjamin Kane, Lenhart Schubert

Many practical applications of dialogue technology require the generation of responses according to a particular developer-specified persona. While a variety of personas can be elicited from recent large language models,…

Dialogue GenerationLanguage ModellingLarge Language Model

SOMA-SQL: Resolving Multi-Source Ambiguity in NL-to-SQL via Synthetic Log and Execution Probing

2026-06-09 · Sai Ashish Somayajula, Marianne Menglin Liu, Chuan Lei, Fjona Parllaku 외 arxiv

Natural language interfaces to databases aim to translate user questions into executable SQL, yet remain brittle in real-world settings where questions are underspecified and schemas are large and ambiguous. Ambiguity ac…

Exploring Database Normalization Effects on SQL Generation

2025-10-02 · Ryosuke Kohita arxiv

Schema design, particularly normalization, is a critical yet often overlooked factor in natural language to SQL (NL2SQL) systems. Most prior research evaluates models on fixed schemas, overlooking the influence of design…

Type prediction