paper-with-me

홈 › Papers

NuclearQA: A Human-Made Benchmark for Language Models for the Nuclear Domain

2023-10-17 · Anurag Acharya, Sai Munikoti, Aaron Hellinger, Sara Smith, Sridevi Wagle, Sameera Horawalavithana

As LLMs have become increasingly popular, they have been used in almost every field. But as the application for LLMs expands from generic fields to narrow, focused science domains, there exists an ever-increasing gap in ways to evaluate their efficacy in those fields. For the benchmarks that do exist, a lot of them focus on questions that don't require proper understanding of the subject in question. In this paper, we present NuclearQA, a human-made benchmark of 100 questions to evaluate language models in the nuclear domain, consisting of a varying collection of questions that have been specifically designed by experts to test the abilities of language models. We detail our approach and show how the mix of several types of questions makes our benchmark uniquely capable of evaluating models in the nuclear domain. We also present our own evaluation metric for assessing LLM's performances due to the limitations of existing ones. Our experiments on state-of-the-art models suggest that even the best LLMs perform less than satisfactorily on our benchmark, demonstrating the scientific knowledge gap of existing LLMs.

📄 PDF Abstract BibTeX arXiv:2310.10920

Code (1)

pnnl/expert2 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models

2026-06-25 · Henry Shaowu Yuchi, Michal Kucer, Benjamin H. Sims, Selma Peterson 외 arxiv

Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant challenge. In nuclear engineering, problem …

Question Generation

NukeBERT: A Pre-trained language model for Low Resource Nuclear Domain

2020-03-30 · Ayush Jain, Dr. N. M. Meenachi, Dr. B. Venkatraman

Significant advances have been made in recent years on Natural Language Processing with machines surpassing human performance in many tasks, including but not limited to Question Answering. The majority of deep learning …

Language ModelingLanguage ModellingQuestion Answering

Myotubularin MTM1 Involved in Centronuclear Myopathy and its Roles in Human and Yeast Cells

2018-04-23

Mutations in the MTM1 gene, encoding the phosphoinositide phosphatase myotubularin, are responsible for the X-linked centronuclear myopathy (XLCNM) or X-linked myotubular myopathy (XLMTM). The MTM1 gene was first identif…

Specificity

Instance Migration Diffusion for Nuclear Instance Segmentation in Pathology

2025-04-02 · Lirui Qi, Hongliang He, Tong Wang, Siwei Feng 외

Nuclear instance segmentation plays a vital role in disease diagnosis within digital pathology. However, limited labeled data in pathological images restricts the overall performance of nuclear instance segmentation. To …

Data AugmentationInstance SegmentationSegmentationSemantic Segmentation

Do Nuclear Submarines Have Nuclear Captains? A Challenge Dataset for Commonsense Reasoning over Adjectives and Objects

2019-11-01 · IJCNLP 2019 11 · James Mullenbach, Jonathan Gordon, Nanyun Peng, Jonathan May

How do adjectives project from a noun to its parts? If a motorcycle is red, are its wheels red? Is a nuclear submarine{'}s captain nuclear? These questions are easy for humans to judge using our commonsense understanding…

Word Embeddings