paper-with-me

Papers

ByteScience: Bridging Unstructured Scientific Literature and Structured Data with Auto Fine-tuned Large Language Model in Token Granularity

2024-11-18 · Tong Xie, Hanzhi Zhang, Shaozhou Wang, Yuwei Wan, Imran Razzak, Chunyu Kit, Wenjie Zhang, Bram Hoex

Natural Language Processing (NLP) is widely used to supply summarization ability from long context to structured information. However, extracting structured knowledge from scientific text by NLP models remains a challenge because of its domain-specific nature to complex data preprocessing and the granularity of multi-layered device-level information. To address this, we introduce ByteScience, a non-profit cloud-based auto fine-tuned Large Language Model (LLM) platform, which is designed to extract structured scientific data and synthesize new scientific knowledge from vast scientific corpora. The platform capitalizes on DARWIN, an open-source, fine-tuned LLM dedicated to natural science. The platform was built on Amazon Web Services (AWS) and provides an automated, user-friendly workflow for custom model development and data extraction. The platform achieves remarkable accuracy with only a small amount of well-annotated articles. This innovative tool streamlines the transition from the science literature to structured knowledge and data and benefits the advancements in natural informatics.

📄 PDF Abstract BibTeX arXiv:2411.12000

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesLanguage ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

SpectraQuery: A Hybrid Retrieval-Augmented Conversational Assistant for Battery Science

2026-01-14 · Sreya Vangara, Jagjit Nanda, Yan-Kai Tzeng, Eric Darve arxiv

Scientific reasoning increasingly requires linking structured experimental data with the unstructured literature that explains it, yet most large language model (LLM) assistants cannot reason jointly across these modalit…

Semantic Parsing

What's In Your Field? Mapping Scientific Research with Knowledge Graphs and Large Language Models

2025-03-12 · Abhipsha Das, Nicholas Lourie, Siavash Golkar, Mariel Pettee

The scientific literature's exponential growth makes it increasingly challenging to navigate and synthesize knowledge across disciplines. Large language models (LLMs) are powerful tools for understanding scientific text,…

Knowledge GraphsNavigateRetrieval-augmented Generation

PubSqueezer: A Text-Mining Web Tool to Transform Unstructured Documents into Structured Data

2020-11-05 · Alberto Calderone

The amount of scientific papers published every day is daunting and constantly increasing. Keeping up with literature represents a challenge. If one wants to start exploring new topics it is hard to have a big picture wi…

ArticlesSentence

Enhancing Scientific Literature Chatbots with Retrieval-Augmented Generation: A Performance Evaluation of Vector and Graph-Based Systems

2026-02-19 · Hamideh Ghanadian, Amin Kamali, Mohammad Hossein Tekieh arxiv

This paper investigates the enhancement of scientific literature chatbots through retrieval-augmented generation (RAG), with a focus on evaluating vector- and graph-based retrieval systems. The proposed chatbot leverages…

Decision Making

LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature

2025-10-28 · Magdalena Lederbauer, Siddharth Betala, Xiyao Li, Ayush Jain 외 arxiv

The development of synthesis procedures remains a fundamental challenge in materials discovery, with procedural knowledge scattered across decades of scientific literature in unstructured formats that are challenging for…