paper-with-me

Papers

OntoURL: A Benchmark for Evaluating Large Language Models on Symbolic Ontological Understanding, Reasoning and Learning

2025-05-16 · Xiao Zhang, Huiyuan Lai, Qianru Meng, Johan Bos

Large language models (LLMs) have demonstrated remarkable capabilities across a range of natural language processing tasks, yet their ability to process structured symbolic knowledge remains underexplored. To address this gap, we propose a taxonomy of LLMs' ontological capabilities and introduce OntoURL, the first comprehensive benchmark designed to systematically evaluate LLMs' proficiency in handling ontologies -- formal, symbolic representations of domain knowledge through concepts, relationships, and instances. Based on the proposed taxonomy, OntoURL systematically assesses three dimensions: understanding, reasoning, and learning through 15 distinct tasks comprising 58,981 questions derived from 40 ontologies across 8 domains. Experiments with 20 open-source LLMs reveal significant performance differences across models, tasks, and domains, with current LLMs showing proficiency in understanding ontological knowledge but substantial weaknesses in reasoning and learning tasks. These findings highlight fundamental limitations in LLMs' capability to process symbolic knowledge and establish OntoURL as a critical benchmark for advancing the integration of LLMs with formal knowledge representations.

📄 PDF Abstract BibTeX arXiv:2505.11031

Code (2)

lastdance500/bench_construct 공식 구현
lastdance500/ontourl 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Weighted Contourlet Parametric (WCP) Feature Based Breast Tumor Classification from B-Mode Ultrasound Image

2021-02-11 · Shahriar Mahmud Kabir, Md. Sayed Tanveer, ASM Shihavuddin, Mohammed Imamul Hassan Bhuiyan

Automated detection of breast tumor in early stages using B-Mode Ultrasound image is crucial for preventing widespread breast cancer specially among women. This paper is primarily focusing on the classification of breast…

General Classification

BaziQA-Benchmark: Evaluating Symbolic and Temporally Compositional Reasoning in Large Language Models

2026-02-13 · Jiangxi Chen, Qian Liu arxiv

We present BaziQA-Benchmark, a standardized benchmark for evaluating symbolic and temporally compositional reasoning in large language models. The benchmark is derived from 200 professionally curated, multiple-choice pro…

Neural Contourlet Network for Monocular 360 Depth Estimation

2022-08-03 · Zhijie Shen, Chunyu Lin, Lang Nie, Kang Liao 외

For a monocular 360 image, depth estimation is a challenging because the distortion increases along the latitude. To perceive the distortion, existing methods devote to designing a deep and complex network architecture. …

DecoderDepth Estimation

A New Multifocus Image Fusion Method Using Contourlet Transform

2017-09-13 · Fatemeh Vakili Moghadam, Hamid Reza Shahdoosti

A new multifocus image fusion approach is presented in this paper. First the contourlet transform is used to decompose the source images into different components. Then, some salient features are extracted from component…

Séparation en composantes structures, textures et bruit d'une image, apport de l'utilisation des contourlettes

2024-11-11 · Jerome Gilles

In this paper, we propose to improve image decomposition algorithms in the case of noisy images. In \cite{gilles1,aujoluvw}, the authors propose to separate structures, textures and noise from an image. Unfortunately, th…