paper-with-me

홈 › Papers

Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking

2026-08-18 · Swati Rajwal, Sanjay Das, Tirthankar Ghosal arxiv

Large language models (LLMs) are increasingly used for scientific hypothesis generation. However, evaluating generated hypotheses remains a challenge for trustworthy AI-enabled scientific workflows. Existing approaches often use LLMs as judges or rely on semantic similarity, which can favor familiar ideas over novel ones. We propose a logit-based energy scoring method that evaluates hypotheses using a language model's intrinsic confidence rather than comparative judgment. We benchmarked seven language models on 1,323 papers across 12 disciplines. Each paper was paired with its hypothesis and fifteen incorrect alternatives. Intrinsic scoring reached 33.0% Hit@1 pooled across both scorers, compared with 16.6% for prompted listwise ranking. The strongest configuration, a 1-billion-parameter model using logit-based energy scoring, reached 53.1%, though this was the maximum across 14 model-by-scorer combinations selected post hoc. Overall, intrinsic model confidence shows potential for scientific hypothesis evaluation. This study also motivates future research on confidence-based methods for trustworthy AI-enabled scientific discovery.

📄 PDF Abstract BibTeX arXiv:2608.17270

Code (1)

Tavish9/awesome-daily-AI-arxiv ★ 113

Tasks

Semantic Similarity

Similar Papers 제목 키워드 기반

Enhancing English Writing Proficiency in China's Polytechnic Students An In-Depth Literature Review on the Application of the Input Hypothesis

2023-11-04 · Wei Zhou

Having good English writing skills is extremely important for students in polytechnic institutions. However, a lot of students in technical schools have difficulties in reaching high levels of skill. The Input Hypothesis…

Enhancing Multi-hop Reasoning through Knowledge Erasure in Large Language Model Editing

2024-08-22 · Mengqi Zhang, Bowen Fang, Qiang Liu, Pengjie Ren 외

Large language models (LLMs) face challenges with internal knowledge inaccuracies and outdated information. Knowledge editing has emerged as a pivotal approach to mitigate these issues. Although current knowledge editing…

knowledge editingLanguage ModelingLanguage ModellingLarge Language Model+1

Large Language Models Must Be Taught to Know What They Don't Know

2024-06-12 · Sanyam Kapoor, Nate Gruver, Manley Roberts, Katherine Collins 외

When using large language models (LLMs) in high-stakes applications, we need to know when we can trust their predictions. Some works argue that prompting high-performance LLMs is sufficient to produce calibrated uncertai…

The Rediscovery Hypothesis: Language Models Need to Meet Linguistics

2021-03-02 · Vassilina Nikoulina, Maxat Tezekbayev, Nuradil Kozhakhmet, Madina Babazhanova 외

There is an ongoing debate in the NLP community whether modern language models contain linguistic knowledge, recovered through so-called probes. In this paper, we study whether linguistic knowledge is a necessary conditi…

Language ModelingLanguage Modelling

HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation

2025-04-15 · Haokun Liu, Sicong Huang, Jingyu Hu, Yangqiaoyu Zhou 외

There is growing interest in hypothesis generation with large language models (LLMs). However, fundamental questions remain: what makes a good hypothesis, and how can we systematically evaluate methods for hypothesis gen…

Benchmarkingscientific discovery