paper-with-me

Papers

AbsPyramid: Benchmarking the Abstraction Ability of Language Models with a Unified Entailment Graph

2023-11-15 · Zhaowei Wang, Haochen Shi, Weiqi Wang, Tianqing Fang, Hongming Zhang, Sehyun Choi, Xin Liu, Yangqiu Song

Cognitive research indicates that abstraction ability is essential in human intelligence, which remains under-explored in language models. In this paper, we present AbsPyramid, a unified entailment graph of 221K textual descriptions of abstraction knowledge. While existing resources only touch nouns or verbs within simplified events or specific domains, AbsPyramid collects abstract knowledge for three components of diverse events to comprehensively evaluate the abstraction ability of language models in the open domain. Experimental results demonstrate that current LLMs face challenges comprehending abstraction knowledge in zero-shot and few-shot settings. By training on our rich abstraction knowledge, we find LLMs can acquire basic abstraction abilities and generalize to unseen events. In the meantime, we empirically show that our benchmark is comprehensive to enhance LLMs across two previous abstraction tasks.

📄 PDF Abstract BibTeX arXiv:2311.09174

Code (1)

hkust-knowcomp/abspyramid 공식 구현 pytorch

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

ConTSG-Bench: A Unified Benchmark for Conditional Time Series Generation

2026-03-05 · Shaocheng Lan, Shuqi Gu, Zhangzhi Xiong, Kan Ren arxiv

Conditional time series generation plays a critical role in addressing data scarcity and enabling causal analysis in real-world applications. Despite its increasing importance, the field lacks a standardized and systemat…

BenchCouncil's View on Benchmarking AI and Other Emerging Workloads

2019-12-02 · Jianfeng Zhan, Lei Wang, Wanling Gao, Rui Ren

This paper outlines BenchCouncil's view on the challenges, rules, and vision of benchmarking modern workloads like Big Data, AI or machine learning, and Internet Services. We conclude the challenges of benchmarking moder…

Benchmarking

Protein-SE(3): Benchmarking SE(3)-based Generative Models for Protein Structure Design

2025-07-27 · Lang Yu, Zhangyang Gao, Cheng Tan, Qin Chen 외 arxiv

SE(3)-based generative models have shown great promise in protein geometry modeling and effective structure design. However, the field currently lacks a modularized benchmark to enable comprehensive investigation and fai…

Goal-Driven Sequential Data Abstraction

2019-07-29 · ICCV 2019 10 · Umar Riaz Muhammad, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 외

Automatic data abstraction is an important capability for both benchmarking machine intelligence and supporting summarization applications. In the former one asks whether a machine can `understand' enough about the meani…

BenchmarkingGeneral Reinforcement Learningreinforcement-learningReinforcement Learning+1

Bencher: Simple and Reproducible Benchmarking for Black-Box Optimization

2025-05-27 · Leonard Papenmeier, Luigi Nardi

We present Bencher, a modular benchmarking framework for black-box optimization that fundamentally decouples benchmark execution from optimization logic. Unlike prior suites that focus on combining many benchmarks in a s…

Benchmarking