paper-with-me

Papers

Developing a Scalable Benchmark for Assessing Large Language Models in Knowledge Graph Engineering

2023-08-31 · Lars-Peter Meyer, Johannes Frey, Kurt Junghanns, Felix Brei, Kirill Bulert, Sabine Gründer-Fahrer, Michael Martin

As the field of Large Language Models (LLMs) evolves at an accelerated pace, the critical need to assess and monitor their performance emerges. We introduce a benchmarking framework focused on knowledge graph engineering (KGE) accompanied by three challenges addressing syntax and error correction, facts extraction and dataset generation. We show that while being a useful tool, LLMs are yet unfit to assist in knowledge graph generation with zero-shot prompting. Consequently, our LLM-KG-Bench framework provides automatic evaluation and storage of LLM responses as well as statistical data and visualization tools to support tracking of prompt engineering and model performance.

📄 PDF Abstract BibTeX arXiv:2308.16622

Code (2)

aksw/llm-kg-bench 공식 구현
aksw/llm-kg-bench-results 공식 구현

Tasks

BenchmarkingDataset GenerationGraph GenerationKnowledge Graphs

Similar Papers 제목 키워드 기반

CREW-WILDFIRE: Benchmarking Agentic Multi-Agent Collaborations at Scale

2025-07-07 · Jonathan Hyun, Nicholas R Waytowich, Boyuan Chen arxiv

Despite rapid progress in large language model (LLM)-based multi-agent systems, current benchmarks fall short in evaluating their scalability, robustness, and coordination capabilities in complex, dynamic, real-world tas…

Spatial Reasoning

ParaCook: On Time-Efficient Planning for Multi-Agent Systems

2025-10-13 · Shiqi Zhang, Xinbei Ma, Yunqing Xu, Zouying Cao 외 arxiv

Large Language Models (LLMs) exhibit strong reasoning abilities for planning long-horizon, real-world tasks, yet existing agent benchmarks focus on task completion while neglecting time efficiency in parallel and asynchr…

NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models

2026-06-25 · Henry Shaowu Yuchi, Michal Kucer, Benjamin H. Sims, Selma Peterson 외 arxiv

Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant challenge. In nuclear engineering, problem …

Question Generation

Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models

2024-10-11 · Yeeun Kim, Young Rok Choi, Eunkyung Choi, Jinhwan Choi 외

Large language models (LLMs) have demonstrated remarkable performance in the legal domain, with GPT-4 even passing the Uniform Bar Exam in the U.S. However their efficacy remains limited for non-standardized tasks and ta…

Legal ReasoningRAGRetrieval-augmented Generation

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

2025-11-03 · Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner 외 arxiv

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as 'safety'…