paper-with-me

홈 › Papers

ScholarBench: A Bilingual Benchmark for Abstraction, Comprehension, and Reasoning Evaluation in Academic Contexts

2025-05-22 · Dongwon Noh, Donghyeok Koh, Junghun Yuk, Gyuwan Kim, Jaeyong Lee, Kyungtae Lim, Cheoneum Park

Prior benchmarks for evaluating the domain-specific knowledge of large language models (LLMs) lack the scalability to handle complex academic tasks. To address this, we introduce \texttt{ScholarBench}, a benchmark centered on deep expert knowledge and complex academic problem-solving, which evaluates the academic reasoning ability of LLMs and is constructed through a three-step process. \texttt{ScholarBench} targets more specialized and logically complex contexts derived from academic literature, encompassing five distinct problem types. Unlike prior benchmarks, \texttt{ScholarBench} evaluates the abstraction, comprehension, and reasoning capabilities of LLMs across eight distinct research domains. To ensure high-quality evaluation data, we define category-specific example attributes and design questions that are aligned with the characteristic research methodologies and discourse structures of each domain. Additionally, this benchmark operates as an English-Korean bilingual dataset, facilitating simultaneous evaluation for linguistic capabilities of LLMs in both languages. The benchmark comprises 5,031 examples in Korean and 5,309 in English, with even state-of-the-art models like o3-mini achieving an average evaluation score of only 0.543, demonstrating the challenging nature of this benchmark.

📄 PDF Abstract BibTeX arXiv:2505.16566

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Whose Name Comes Up? II: Benchmarking and Intervention-Based Auditing of LLM-Based Scholar Recommendation

2026-02-09 · Lisette Espín-Noboa, Gonzalo Gabriel Méndez arxiv

Large language models (LLMs) are now used for academic expert recommendation. Existing audits typically evaluate such recommendations in isolation, ignoring end-user inference-time interventions. Thus, it remains unclear…

CodeApex: A Bilingual Programming Evaluation Benchmark for Large Language Models

2023-09-05 · Lingyue Fu, Huacan Chai, Shuang Luo, Kounianhua Du 외

With the emergence of Large Language Models (LLMs), there has been a significant improvement in the programming capabilities of models, attracting growing attention from researchers. Evaluating the programming capabiliti…

Code GenerationMultiple-choice

BiPaR: A Bilingual Parallel Dataset for Multilingual and Cross-lingual Reading Comprehension on Novels

2019-10-11 · IJCNLP 2019 11 · Yimin Jing, Deyi Xiong, Yan Zhen

This paper presents BiPaR, a bilingual parallel novel-style machine reading comprehension (MRC) dataset, developed to support multilingual and cross-lingual reading comprehension. The biggest difference between BiPaR and…

coreference-resolutionCoreference ResolutionMachine Reading ComprehensionReading Comprehension+1

LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding

2025-10-19 · Sheikh Jubair, Arwa Omayrah, Amal Alshammari, Alhanoof Althnian 외 arxiv

Recent advancements in Large Language Models (LLMs) have demonstrated sophisticated capabilities, including the ability to process and comprehend extended contexts. These emergent capabilities necessitate rigorous evalua…

Long-Context UnderstandingInformation ExtractionQuestion Answering

BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

2026-08-18 · Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov 외 arxiv

While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks e…

Information Extraction