paper-with-me

홈 › Papers

ELAIPBench: A Benchmark for Expert-Level Artificial Intelligence Paper Understanding

2025-10-12 · Xinbang Dai, Huikang Hu, Yongrui Chen, Jiaqi Li, Rihui Jin, Yuyang Zhang, Xiaoguang Li, Lifeng Shang, Guilin Qi arxiv

While large language models (LLMs) excel at many domain-specific tasks, their ability to deeply comprehend and reason about full-length academic papers remains underexplored. Existing benchmarks often fall short of capturing such depth, either due to surface-level question design or unreliable evaluation metrics. To address this gap, we introduce ELAIPBench, a benchmark curated by domain experts to evaluate LLMs' comprehension of artificial intelligence (AI) research papers. Developed through an incentive-driven, adversarial annotation process, ELAIPBench features 403 multiple-choice questions from 137 papers. It spans three difficulty levels and emphasizes non-trivial reasoning rather than shallow retrieval. Our experiments show that the best-performing LLM achieves an accuracy of only 39.95%, far below human performance. Moreover, we observe that frontier LLMs equipped with a thinking mode or a retrieval-augmented generation (RAG) system fail to improve final results-even harming accuracy due to overthinking or noisy retrieval. These findings underscore the significant gap between current LLM capabilities and genuine comprehension of academic papers.

📄 PDF Abstract BibTeX arXiv:2510.10549

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence

2025-12-04 · Wenjin Liu, Haoran Luo, Xin Feng, Xiang Ji 외 arxiv

Legal general intelligence (GI) refers to artificial intelligence (AI) that encompasses legal understanding, reasoning, and decision-making, simulating the expertise of legal experts across domains. However, existing ben…

Artificial General Intelligence, Existential Risk, and Human Risk Perception

2023-11-15 · David R. Mandel

Artificial general intelligence (AGI) does not yet exist, but given the pace of technological development in artificial intelligence, it is projected to reach human-level intelligence within roughly the next two decades.…

Future progress in artificial intelligence: A survey of expert opinion

2025-08-09 · Vincent C. Müller, Nick Bostrom arxiv

There is, in some quarters, concern about high-level machine intelligence and superintelligent AI coming up in a few decades, bringing with it significant risks for humanity. In other quarters, these issues are ignored o…

Bridging the Gap between Artificial Intelligence and Artificial General Intelligence: A Ten Commandment Framework for Human-Like Intelligence

2022-10-17 · Ananta Nair, Farnoush Banaei-Kashani

The field of artificial intelligence has seen explosive growth and exponential success. The last phase of development showcased deep learnings ability to solve a variety of difficult problems across a multitude of domain…

ASI-Bench: At the Dawn of Artificial Superintelligence

2026-08-18 · Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou 외 arxiv

Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of…