paper-with-me

홈 › Papers

LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning

2025-09-24 · You-Le Fang, Dong-Shan Jian, Xiang Li, Ce Meng, Ling-Shi Meng, Chen-Xu Yan, Zhi-Zhang Bian, Yan-Qing Ma arxiv

While Large Language Models (LLMs) excel in general domains, their reliability often falls short in scientific problem-solving. The advancement of scientific AI depends on large-scale, high-quality corpora. However, existing scientific question-answering (QA) datasets suffer from high error rates, frequently resulting from logical leaps and implicit reasoning within the answers. To address this issue, we introduce LOCA (Logical Chain Augmentation), a novel framework for automatically cleaning scientific corpora, implemented through an augment-and-review loop. At its core, LOCA enhances raw answers by completing missing logical steps and explicitly separating the underlying scientific principle from its subsequent derivation. By applying LOCA to challenging scientific corpora, we demonstrate that it can automatically filter noisy datasets, typically reducing the error rate from as high as 20\% to below 2\%. LOCA provides a scalable and effective methodology for creating high-quality scientific corpora, paving the way for more reliable training and evaluation of scientific AI.

📄 PDF Abstract BibTeX arXiv:2510.01249

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RAISE: Enhancing Scientific Reasoning in LLMs via Step-by-Step Retrieval

2025-06-10 · Minhae Oh, Jeonghye Kim, Nakyung Lee, Donggeon Seo 외

Scientific reasoning requires not only long-chain reasoning processes, but also knowledge of domain-specific terminologies and adaptation to updated findings. To deal with these challenges for scientific reasoning, we in…

Problem DecompositionRetrieval

TDMSci: A Specialized Corpus for Scientific Literature Entity Tagging of Tasks Datasets and Metrics

2021-01-25 · EACL 2021 2 · Yufang Hou, Charles Jochim, Martin Gleize, Francesca Bonin 외

Tasks, Datasets and Evaluation Metrics are important concepts for understanding experimental scientific papers. However, most previous work on information extraction for scientific literature mainly focuses on the abstra…

Data Augmentation

Corpus Augmentation by Sentence Segmentation for Low-Resource Neural Machine Translation

2019-05-22 · Jinyi Zhang, Tadahiro Matsumoto

Neural Machine Translation (NMT) has been proven to achieve impressive results. The NMT system translation results depend strongly on the size and quality of parallel corpora. Nevertheless, for many language pairs, no ri…

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationNMT+3

Transfer Learning for Scientific Data Chain Extraction in Small Chemical Corpus with BERT-CRF Model

2019-05-13 · Na Pang, Li Qian, Weimin Lyu, Jin-Dong Yang

Computational chemistry develops fast in recent years due to the rapid growth and breakthroughs in AI. Thanks for the progress in natural language processing, researchers can extract more fine-grained knowledge in public…

Computational chemistryEntity Extraction using GANNERTransfer Learning

Modeling Changing Scientific Concepts with Complex Networks: A Case Study on the Chemical Revolution

2026-03-18 · Sofía Aguilar-Valdez, Stefania Degaetano-Ortlieb arxiv

While context embeddings produced by LLMs can be used to estimate conceptual change, these representations are often not interpretable nor time-aware. Moreover, bias augmentation in historical data poses a non-trivial ri…