paper-with-me

홈 › Papers

Does Prompt Design Impact Quality of Data Imputation by LLMs?

2025-06-04 · Shreenidhi Srinivasan, Lydia Manikonda

Generating realistic synthetic tabular data presents a critical challenge in machine learning. It adds another layer of complexity when this data contain class imbalance problems. This paper presents a novel token-aware data imputation method that leverages the in-context learning capabilities of large language models. This is achieved through the combination of a structured group-wise CSV-style prompting technique and the elimination of irrelevant contextual information in the input prompt. We test this approach with two class-imbalanced binary classification datasets and evaluate the effectiveness of imputation using classification-based evaluation metrics. The experimental results demonstrate that our approach significantly reduces the input prompt size while maintaining or improving imputation quality compared to our baseline prompt, especially for datasets that are of relatively smaller in size. The contributions of this presented work is two-fold -- 1) it sheds light on the importance of prompt design when leveraging LLMs for synthetic data generation and 2) it addresses a critical gap in LLM-based data imputation for class-imbalanced datasets with missing data by providing a practical solution within computational constraints. We hope that our work will foster further research and discussions about leveraging the incredible potential of LLMs and prompt engineering techniques for synthetic data generation.

📄 PDF Abstract BibTeX arXiv:2506.04172

Code (0)

등록된 구현이 없습니다.

Tasks

Binary ClassificationImputationIn-Context LearningPrompt EngineeringSynthetic Data Generation

Similar Papers 제목 키워드 기반

The Impact of Prompt Programming on Function-Level Code Generation

2024-12-29 · Ranim Khojah, Francisco Gomes de Oliveira Neto, Mazen Mohamad, Philipp Leitner

Large Language Models (LLMs) are increasingly used by software engineers for code generation. However, limitations of LLMs such as irrelevant or incorrect code have highlighted the need for prompt programming (or prompt …

Code GenerationPrompt Engineering

Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance

2024-02-22 · Ziqi Yin, Hao Wang, Kaito Horio, Daisuke Kawahara 외

We investigate the impact of politeness levels in prompts on the performance of large language models (LLMs). Polite language in human communications often garners more compliance and effectiveness, while rudeness can ca…

Does AI Homogenize Student Thinking? A Multi-Dimensional Analysis of Structural Convergence in AI-Augmented Essays

2026-03-22 · Keito Inoshita, Michiaki Omura, Tsukasa Yamanaka, Go Maeda 외 arxiv

While AI-assisted writing has been widely reported to improve essay quality, its impact on the structural diversity of student thinking remains unexplored. Analyzing 6,875 essays across five conditions (Human-only, AI-on…

The Hidden Cost of an Image: Quantifying the Energy Consumption of AI Image Generation

2025-06-20 · Giulia Bertazzini, Chiara Albisani, Daniele Baracchi, Dasara Shullani 외

With the growing adoption of AI image generation, in conjunction with the ever-increasing environmental resources demanded by AI, we are urged to answer a fundamental question: What is the environmental impact hidden beh…

Image GenerationQuantization

Prompt Design Matters for Computational Social Science Tasks but in Unpredictable Ways

2024-06-17 · Shubham Atreja, Joshua Ashkinaze, Lingyao Li, Julia Mendelsohn 외

Manually annotating data for computational social science tasks can be costly, time-consuming, and emotionally draining. While recent work suggests that LLMs can perform such annotation tasks in zero-shot settings, littl…

Model Selection