paper-with-me

홈 › Papers

Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models

2026-04-20 · Seyedali Mohammadi, Manas Gaur, Francis Ferraro arxiv

Scientific feasibility assessment asks whether a claim is consistent with established knowledge and whether experimental evidence could support or refute it. We frame feasibility assessment as a diagnostic reasoning task in which, given a hypothesis, a model predicts feasible or infeasible and justifies its decision. We evaluate large language models (LLMs) under controlled knowledge conditions (hypothesis-only, with experiments, with outcomes, or both) and probe robustness by progressively removing portions of the experimental and/or outcome context. Across multiple LLMs and two datasets, providing outcome evidence is generally more reliable than providing experiment descriptions. Outcomes tend to improve accuracy beyond what internal knowledge alone provides, whereas experimental text can be brittle and may degrade performance when the context is incomplete. These findings clarify when experimental evidence benefits LLM-based feasibility assessment and when it introduces fragility.

📄 PDF Abstract BibTeX arXiv:2604.18786

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows

2025-12-18 · Wanghan Xu, Yuhao Zhou, Yifan Zhou, Qinglong Cao 외 arxiv

Despite advances in scientific AI, a coherent framework for Scientific General Intelligence (SGI)-the ability to autonomously conceive, investigate, and reason across scientific domains-remains lacking. We present an ope…

Reinforcement Learning

Exploring Early Prediction of Buyer-Seller Negotiation Outcomes

2020-04-06 · Kushal Chawla, Gale Lucas, Jonathan May, Jonathan Gratch

Agents that negotiate with humans find broad applications in pedagogy and conversational AI. Most efforts in human-agent negotiations rely on restrictive menu-driven interfaces for communication. To advance the research …

Language ModelingLanguage ModellingPredictionSentence

HARPA: A Testability-Driven, Literature-Grounded Framework for Research Ideation

2025-10-01 · Rosni Vasu, Peter Jansen, Pao Siangliulue, Cristina Sarasua 외 arxiv

While there has been a surge of interest in automated scientific discovery (ASD), especially with the emergence of LLMs, it remains challenging for tools to generate hypotheses that are both testable and grounded in the …

SABER: Stability-Aware Early Exit for LLM Reasoning via Adversarial Branch Probing

2026-08-28 · Wanli Cheng, Haiya Xiang, Juntao Li, Hongling Wang 외 arxiv

Large Reasoning Models (LRMs) achieve strong reasoning capabilities, yet long-chain reasoning becomes inefficient once the intermediate answer stabilizes across reasoning steps: additional reasoning yields little margina…

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?

2026-04-12 · Udari Madhushani Sehwag, Elaine Lau, Haniyeh Ehsani Oskouie, Shayan Shabihi 외 arxiv

Accelerating scientific discovery requires the identification of which experiments would yield the best outcomes before committing resources to costly physical validation. While existing benchmarks evaluate LLMs on scien…