paper-with-me

홈 › Papers

SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection

2026-03-23 · Kexian Tang, Jiani Wang, Shaowen Wang, Kaifeng Lyu arxiv

While large language models (LLMs) are pretrained on massive amounts of data, their knowledge coverage remains incomplete in specialized, data-scarce domains, motivating extensive efforts to study synthetic data generation for knowledge injection. We propose SPA (Scaling Prompt-engineered Augmentation), a simple but tough-to-beat baseline that uses a small set of carefully designed prompts to generate large-scale synthetic data for knowledge injection. Through systematic comparisons, we find that SPA outperforms several strong baselines. Furthermore, we identify two key limitations of prior approaches: (1) while RL-based methods may improve the token efficiency of LLM-based data augmentation at small scale, they suffer from diversity collapse as data scales, leading to diminishing returns; and (2) while multi-stage prompting may outperform simple augmentation methods, their advantages can disappear after careful prompt tuning. Our results suggest that, for knowledge injection, careful prompt design combined with straightforward large-scale augmentation can be surprisingly effective, and we hope SPA can serve as a strong baseline for future studies in this area. Our code is available at https://github.com/Tangkexian/SPA.

📄 PDF Abstract BibTeX arXiv:2603.22213

Code (0)

등록된 구현이 없습니다.

Tasks

Synthetic Data GenerationData Augmentation

Similar Papers 제목 키워드 기반

A simple but tough-to-beat baseline for the Fake News Challenge stance detection task

2017-07-11 · Benjamin Riedel, Isabelle Augenstein, Georgios P. Spithourakis, Sebastian Riedel

Identifying public misinformation is a complicated and challenging task. An important part of checking the veracity of a specific claim is to evaluate the stance different news sources take towards the assertion. Automat…

Fact CheckingMisinformationStance Detection

Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learning

2025-05-16 · Yuanzhao Zhang, William Gilpin

Recently-developed time series foundation models for scientific machine learning exhibit emergent abilities to predict physical systems. These abilities include zero-shot forecasting, in which a model forecasts future st…

In-Context LearningTime SeriesTime Series Forecasting

Cosine meets Softmax: A tough-to-beat baseline for visual grounding

2020-09-13 · Nivedita Rufus, Unni Krishnan R Nair, K. Madhava Krishna, Vineet Gandhi

In this paper, we present a simple baseline for visual grounding for autonomous driving which outperforms the state of the art methods, while retaining minimal design choices. Our framework minimizes the cross-entropy lo…

Autonomous DrivingMetric LearningReferring Expression ComprehensionSentence+1

qDKT: Question-centric Deep Knowledge Tracing

2020-05-25 · Shashank Sonkar, Andrew E. Waters, Andrew S. Lan, Phillip J. Grimaldi 외

Knowledge tracing (KT) models, e.g., the deep knowledge tracing (DKT) model, track an individual learner's acquisition of skills over time by examining the learner's performance on questions related to those skills. A pr…

Knowledge TracingLanguage ModelingLanguage Modelling

simpleKT: A Simple But Tough-to-Beat Baseline for Knowledge Tracing

2023-02-14 · Zitao Liu, Qiongqiong Liu, Jiahao Chen, Shuyan Huang 외

Knowledge tracing (KT) is the problem of predicting students' future performance based on their historical interactions with intelligent tutoring systems. Recently, many works present lots of special methods for applying…

Knowledge Tracing