paper-with-me

홈 › Papers

Fine-tuning for Better Few Shot Prompting: An Empirical Comparison for Short Answer Grading

2025-08-06 · Joel Walsh, Siddarth Mamidanna, Benjamin Nye, Mark Core, Daniel Auerbach arxiv

Research to improve Automated Short Answer Grading has recently focused on Large Language Models (LLMs) with prompt engineering and no- or few-shot prompting to achieve best results. This is in contrast to the fine-tuning approach, which has historically required large-scale compute clusters inaccessible to most users. New closed-model approaches such as OpenAI's fine-tuning service promise results with as few as 100 examples, while methods using open weights such as quantized low-rank adaptive (QLORA) can be used to fine-tune models on consumer GPUs. We evaluate both of these fine-tuning methods, measuring their interaction with few-shot prompting for automated short answer grading (ASAG) with structured (JSON) outputs. Our results show that finetuning with small amounts of data has limited utility for Llama open-weight models, but that fine-tuning methods can outperform few-shot baseline instruction-tuned LLMs for OpenAI's closed models. While our evaluation set is limited, we find some evidence that the observed benefits of finetuning may be impacted by the domain subject matter. Lastly, we observed dramatic improvement with the LLama 3.1 8B-Instruct open-weight model by seeding the initial training examples with a significant amount of cheaply generated synthetic training data.

📄 PDF Abstract BibTeX arXiv:2508.04063

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Similar Papers 제목 키워드 기반

Enhancing Cross-lingual Prompting with Dual Prompt Augmentation

2022-02-15 · Meng Zhou, Xin Li, Yue Jiang, Lidong Bing

Prompting shows promising results in few-shot scenarios. However, its strength for multilingual/cross-lingual problems has not been fully exploited. Zhao and Sch\"utze (2021) made initial explorations in this direction b…

Cross-Lingual Transfer

Revisiting Automated Prompting: Are We Actually Doing Better?

2023-04-07 · Yulin Zhou, Yiren Zhao, Ilia Shumailov, Robert Mullins 외

Current literature demonstrates that Large Language Models (LLMs) are great few-shot learners, and prompting significantly increases their performance on a range of downstream tasks in a few-shot learning setting. An att…

Few-Shot Learning

Towards Unified Prompt Tuning for Few-shot Learning

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Prompt-based fine-tuning has boosted the performance of Pre-trained Language Models (PLMs) on few-shot learning by employing task-specific prompts. However, PLMs are unfamiliar with the prompt-style expressions during pr…

Few-Shot LearningLanguage ModelingLanguage ModellingMasked Language Modeling+2

Discrete and Soft Prompting for Multilingual Models

2021-09-08 · EMNLP 2021 11 · Mengjie Zhao, Hinrich Schütze

It has been shown for English that discrete and soft prompting perform strongly in few-shot learning with pretrained language models (PLMs). In this paper, we show that discrete and soft prompting perform better than fin…

Few-Shot LearningNatural Language Inference

Towards Unified Prompt Tuning for Few-shot Text Classification

2022-05-11 · Jianing Wang, Chengyu Wang, Fuli Luo, Chuanqi Tan 외

Prompt-based fine-tuning has boosted the performance of Pre-trained Language Models (PLMs) on few-shot text classification by employing task-specific prompts. Yet, PLMs are unfamiliar with prompt-style expressions during…

ClassificationFew-Shot LearningFew-Shot Text ClassificationLanguage Modeling+6