paper-with-me

홈 › Papers

OpenBioRQ: Unsolved Biomedical Research Questions for Agents

2026-06-20 · Minbyul Jeong arxiv

A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the claim. I find that current agentic models rarely fabricate citations (over $99\%$ resolve), yet roughly $15.9\%$ link to the wrong paper. Existing benchmarks miss this failure mode: when a question has a fixed answer key, a model can reproduce the expected source from that key rather than independently verifying that the source supports the claim. I introduce \textbf{\openbiorq{}}, a retrieval-grounded agentic benchmark of $12{,}553$ unsolved biomedical research questions across $12$ domains that treats open questions as a faithfulness-and-abstention probe. To my knowledge, this is the first biomedical benchmark to combine an agentic setting -- where the model must issue multiple tool calls -- with unsolved questions that have no answer key. Openness is verified against real follow-up evidence rather than a model's parametric knowledge. Difficulty is empirical: I anchor it on questions that three open-weight reference models fail to answer, rather than on subjective hardness labels. On this hardest subset, held-out models from the same lineage as the difficulty anchors solve only ~17%, while three independent frontier agents (Gemini-3-Pro, Opus-4.7, GPT-5.5) span a wide 29-60% range. The benchmark is thus hard, non-saturating (the best agent still leaves ~33-40\% unsolved), and discriminating across capability tiers. Beyond difficulty, I observe agentic collapse on the hardest questions, where agents stop using their tools. For the most collapse-prone model, blocking tool access entirely barely changes its score -- so tools stop paying off exactly where they are needed most. A frozen per-question checklist raises inter-judge agreement from Spearman 0.35 to 0.82.

📄 PDF Abstract BibTeX arXiv:2606.21959

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Benchmarks in Leipzig

2026-06-04 · Andrei Balakin, Miklós Bóna, Marie-Charlotte Brandenburg, Clara Briand 외 arxiv

Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the 3-day workshop *Benchmarks in Leipzig* wi…

Mathematical Reasoning

UQ: Assessing Language Models on Unsolved Questions

2025-08-25 · Fan Nie, Ken Ziyu Liu, Zihao Wang, Rui Sun 외 arxiv

Benchmarks shape progress in AI research. A useful benchmark should be both difficult and realistic: questions should challenge frontier models while also reflecting real-world usage. Yet, current paradigms face a diffic…

Towards Artificial Intelligence Research Assistant for Expert-Involved Learning

2025-05-03 · Tianyu Liu, Simeng Han, Xiao Luo, Hanchen Wang 외

Large Language Models (LLMs) and Large Multi-Modal Models (LMMs) have emerged as transformative tools in scientific research, yet their reliability and specific contributions to biomedical applications remain insufficien…

ArticlesPrompt Engineering

Biomedical Concept Normalization by Leveraging Hypernyms

2021-11-01 · EMNLP 2021 11 · Cheng Yan, Yuanzhe Zhang, Kang Liu, Jun Zhao 외

Biomedical Concept Normalization (BCN) is widely used in biomedical text processing as a fundamental module. Owing to numerous surface variants of biomedical concepts, BCN still remains challenging and unsolved. In this …

Autonomous self-evolving research on biomedical data: the DREAM paradigm

2024-07-18 · Luojia Deng, Yijie Wu, Yongyong Ren, Hui Lu

In contemporary biomedical research, the efficiency of data-driven approaches is hindered by large data volumes, tool selection complexity, and human resource limitations, necessitating the development of fully autonomou…

Articles