paper-with-me

홈 › Papers

Bridging the Plausibility-Validity Gap by Fine-Tuning a Reasoning-Enhanced LLM for Chemical Synthesis and Discovery

2025-07-09 · Malikussaid, Hilal Hudan Nuha, Isman Kurniawan arxiv

Large Language Models frequently generate outputs that appear scientifically reasonable yet violate fundamental principles--a phenomenon we characterize as the "plausibility-validity gap." This challenge proves especially acute in chemistry, where superficial correctness masks deeper errors in molecular structure, reaction mechanisms, and synthetic pathways. We present a systematic approach combining a reasoning-centric model architecture (Magistral Small) with Low-Rank Adaptation fine-tuning on a dual-domain dataset covering molecular properties and chemical transformations. Evaluation reveals substantial improvements: the fine-tuned system achieves 96.3% format adherence, 97.4% chemical validity, and 74.4% synthesis feasibility. Comparative analysis shows our approach outperforms specialized translation models like MolT5 (97.4% vs 77.2% validity) while achieving performance comparable to complex tool-augmented systems like ChemCrow (9.0/10 vs 9.24/10 expert rating) through a more transparent, efficient methodology. Results demonstrate a learning hierarchy where syntactic correctness develops before chemical understanding, which precedes synthetic planning capability. This work establishes a reproducible framework for transforming generalist language models into dependable scientific tools while identifying critical areas including stereochemical precision, knowledge currency, and computational accessibility as key challenges for future advancement.

📄 PDF Abstract BibTeX arXiv:2507.07328

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects

2025-10-08 · Leonardo Bertolazzi, Sandro Pezzelle, Raffaella Bernardi arxiv

Both humans and large language models (LLMs) exhibit content effects: biases in which the plausibility of the semantic content of a reasoning problem influences judgments regarding its logical validity. While this phenom…

CAREAgent: Clinical Agent with Structured Reasoning and Tool-Integrated for Order Generation

2026-05-31 · Ruihui Hou, Ziyue Huai, Chennuo Zhang, Ziyan Liu 외 arxiv

Clinical order generation serves as a critical bridge between clinical decision-making and real-world practice, translating medical decisions into concrete and executable orders. Existing agents mainly focus on coarse-gr…

Reinforcement Learning

Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering

2025-05-18 · Marco Valentino, Geonhee Kim, Dhairya Dalal, Zhixue Zhao 외

Large language models (LLMs) frequently demonstrate reasoning limitations, often conflating content plausibility (i.e., material inference) with logical validity (i.e., formal inference). This can result in biased infere…

Language ModelingLanguage Modelling

NCL-UoR at SemEval-2026 Task 5: Embedding-Based Methods, Fine-Tuning, and LLMs for Word Sense Plausibility Rating

2026-03-09 · Tong Wu, Thanet Markchom, Huizhi Liang arxiv

Word sense plausibility rating requires predicting the human-perceived plausibility of a given word sense on a 1-5 scale in the context of short narrative stories containing ambiguous homonyms. This paper systematically …

Bridging the Gap Between Plausibility and Admissibility: Constraint-Aware Flow Maps for Dynamic Graph Systems

2026-07-23 · Michael Romei de Socio, Gian Luca Pozzato, Alessio Merlo arxiv

Generative models can support decision-making under uncertainty by producing ensembles of plausible future system trajectories, but statistical plausibility does not ensure structural feasibility. This study investigates…

Trajectory Modeling