paper-with-me

홈 › Papers

Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy

2026-01-28 · Si Chen, Le Huy Khiem, Annalisa Szymanski, Ronald Metoyer, Ting Hua, Nitesh V. Chawla arxiv

Open-ended question answering (QA) evaluates a model's ability to perform contextualized reasoning beyond factual recall. This challenge is especially acute in practice-based domains, where knowledge is procedural and grounded in professional judgment, while most existing LLM benchmarks depend on pre-existing human exam datasets that are often unavailable in such settings. We introduce a framework for automated benchmark generation from expert-authored guidelines informed by Bloom's Taxonomy. It converts expert practices into implicit violation-based scenarios and expands them into auto-graded multiple-choice questions (MCQs) and multi-turn dialogues across four cognitive levels, enabling deterministic, reproducible, and scalable evaluation. Applied to three applied domains: teaching, dietetics, and caregiving, we find differences between model and human-like reasoning: LLMs sometimes perform relatively better on higher-order reasoning (Analyze) but fail more frequently on lower-level items (Remember). We produce large-scale, psychometrically informed benchmarks that surface these non-intuitive model behaviors and enable evaluation of contextualized reasoning in real-world settings.

📄 PDF Abstract BibTeX arXiv:2601.20253

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Automated Conversion of Music Videos into Lyric Videos

2023-08-28 · Jiaju Ma, Anyi Rao, Li-Yi Wei, Rubaiat Habib Kazi 외

Musicians and fans often produce lyric videos, a form of music videos that showcase the song's lyrics, for their favorite songs. However, making such videos can be challenging and time-consuming as the lyrics need to be …

AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents

2024-03-13 · Yao Fu, Dong-Ki Kim, Jaekyeom Kim, Sungryull Sohn 외

Recent advances in large language models (LLMs) have empowered AI agents capable of performing various sequential decision-making tasks. However, effectively guiding LLMs to perform well in unfamiliar domains like web na…

Decision MakingIn-Context LearningLanguage ModelingLanguage Modelling+2

Automatic Legal Writing Evaluation of LLMs

2025-04-29 · Ramon Pires, Roseval Malaquias Junior, Rodrigo Nogueira

Despite the recent advances in Large Language Models, benchmarks for evaluating legal writing remain scarce due to the inherent complexity of assessing open-ended responses in this domain. One of the key challenges in ev…

Bridging Domain Gaps for Fine-Grained Moth Classification Through Expert-Informed Adaptation and Foundation Model Priors

2025-08-27 · Ross J Gardiner, Guillaume Mougeot, Sareh Rowlands, Benno I Simmons 외 arxiv

Labelling images of Lepidoptera (moths) from automated camera systems is vital for understanding insect declines. However, accurate species identification is challenging due to domain shifts between curated images and no…

Knowledge Distillation

Knowledge-Informed Auto-Penetration Testing Based on Reinforcement Learning with Reward Machine

2024-05-24 · Yuanliang Li, Hanzheng Dai, Jun Yan

Automated penetration testing (AutoPT) based on reinforcement learning (RL) has proven its ability to improve the efficiency of vulnerability identification in information systems. However, RL-based PT encounters several…

Q-LearningReinforcement Learning (RL)