paper-with-me

홈 › Papers

GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation

2025-05-26 · Zihong Chen, Wanli Jiang, Jinzhe Li, Zhonghang Yuan, Huanjun Kong, Wanli Ouyang, Nanqing Dong

Fine-tuning for large language models (LLMs) typically requires substantial amounts of high-quality supervised data, which is both costly and labor-intensive to acquire. While synthetic data generation has emerged as a promising solution, existing approaches frequently suffer from factual inaccuracies, insufficient long-tail coverage, simplistic knowledge structures, and homogenized outputs. To address these challenges, we introduce GraphGen, a knowledge graph-guided framework designed for three key question-answering (QA) scenarios: atomic QA, aggregated QA, and multi-hop QA. It begins by constructing a fine-grained knowledge graph from the source text. It then identifies knowledge gaps in LLMs using the expected calibration error metric, prioritizing the generation of QA pairs that target high-value, long-tail knowledge. Furthermore, GraphGen incorporates multi-hop neighborhood sampling to capture complex relational information and employs style-controlled generation to diversify the resulting QA data. Experimental results on knowledge-intensive tasks under closed-book settings demonstrate that GraphGen outperforms conventional synthetic data methods, offering a more reliable and comprehensive solution to the data scarcity challenge in supervised fine-tuning. The code and data are publicly available at https://github.com/open-sciencelab/GraphGen.

📄 PDF Abstract BibTeX arXiv:2505.20416

Code (1)

open-sciencelab/graphgen 공식 구현

Tasks

Question AnsweringSynthetic Data Generation

Similar Papers 제목 키워드 기반

Are LLMs Effective Backbones for Fine-tuning? An Experimental Investigation of Supervised LLMs on Chinese Short Text Matching

2024-03-29 · Shulin Liu, Chengcheng Xu, Hao liu, TingHao Yu 외

The recent success of Large Language Models (LLMs) has garnered significant attention in both academia and industry. Prior research on LLMs has primarily focused on enhancing or leveraging their generalization capabiliti…

Natural Language UnderstandingText Matching

GraphGen-Redux: a Fast and Lightweight Recurrent Model for labeled Graph Generation

2021-07-18 · Marco Podda, Davide Bacciu

The problem of labeled graph generation is gaining attention in the Deep Learning community. The task is challenging due to the sparse and discrete nature of graph spaces. Several approaches have been proposed in the lit…

Graph Generation

GraphGen+: Advancing Distributed Subgraph Generation and Graph Learning On Industrial Graphs

2025-03-08 · Yue Jin, Yongchao Liu, Chuntao Hong

Graph-based computations are crucial in a wide range of applications, where graphs can scale to trillions of edges. To enable efficient training on such large graphs, mini-batch subgraph sampling is commonly used, which …

Graph Learning

GraphGen: A Scalable Approach to Domain-agnostic Labeled Graph Generation

2020-01-22 · Nikhil Goyal, Harsh Vardhan Jain, Sayan Ranu

Graph generative models have been extensively studied in the data mining literature. While traditional techniques are based on generating structures that adhere to a pre-decided distribution, recent techniques have shift…

Graph Generation

OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference

2025-02-25 · Xiangyu Zhao, Shengyuan Ding, ZiCheng Zhang, Haian Huang 외

Recent advancements in open-source multi-modal large language models (MLLMs) have primarily focused on enhancing foundational capabilities, leaving a significant gap in human preference alignment. This paper introduces O…

Visual Question Answering (VQA)