paper-with-me

홈 › Papers

BLAST: Benchmarking LLMs with ASP-based Structured Testing

2026-04-24 · Manuel Alejandro Borroto Santana, Erica Coppolillo, Francesco Calimeri, Giuseppe Manco, Simona Perri, Francesco Ricca arxiv

Large Language Models (LLMs) have demonstrated remarkable performance across a broad spectrum of tasks, including natural language understanding, dialogue systems, and code generation. Despite evident progress, less attention has been paid to their effectiveness in handling declarative paradigms such as Answer Set Programming (ASP), to date. In this paper we introduce BLAST: The first dedicated benchmarking methodology and associated dataset for evaluating the accuracy of LLMs in generating ASP code. BLAST provides a structured evaluation framework featuring two novel semantic metrics tailored to ASP code generation. The paper presents the results of an empirical evaluation involving ten well-established graph-related problems from the ASP literature and a diverse set of eight state-of-the-art LLMs.

📄 PDF Abstract BibTeX arXiv:2604.22306

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language UnderstandingCode Generation

Similar Papers 제목 키워드 기반

BLAST: Block-Level Adaptive Structured Matrices for Efficient Deep Neural Network Inference

2024-10-28 · Changwoo Lee, Soo Min Kwon, Qing Qu, Hun-Seok Kim

Large-scale foundation models have demonstrated exceptional performance in language and vision tasks. However, the numerous dense matrix-vector operations involved in these large networks pose significant computational c…

StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs

2024-12-23 · Hailin Chen, Fangkai Jiao, Mathieu Ravaut, Nawshad Farruque 외

The rapid development of large language models (LLMs) necessitates robust, unbiased, and scalable methods for evaluating their capabilities. However, human annotations are expensive to scale, model-based evaluations are …

BenchmarkingLogical ReasoningMath

LLMStructBench: Benchmarking Large Language Model Structured Data Extraction

2026-02-16 · Sönke Tenckhoff, Mario Koddenbrock, Erik Rodner arxiv

We present LLMStructBench, a novel benchmark for evaluating Large Language Models (LLMs) on extracting structured data and generating valid JavaScript Object Notation (JSON) outputs from natural-language text. Our open d…

Deep neuroevolution for limited, heterogeneous data: proof-of-concept application to Neuroblastoma brain metastasis using a small virtual pooled image collection

2022-11-26 · Subhanik Purkayastha, Hrithwik Shalu, David Gutman, Shakeel Modak 외

Artificial intelligence (AI) in radiology has made great strides in recent years, but many hurdles remain. Overfitting and lack of generalizability represent important ongoing challenges hindering accurate and dependable…

PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design

2025-12-16 · Ruozhao Yang, Mingfei Cheng, Gelei Deng, Tianwei Zhang 외 arxiv

Penetration testing is essential for assessing and strengthening system security against real-world threats, yet traditional workflows remain highly manual, expertise-intensive, and difficult to scale. Although recent ad…

Domain Adaptation