paper-with-me

Papers

TAGAL: Tabular Data Generation using Agentic LLM Methods

2025-09-04 · Benoît Ronval, Pierre Dupont, Siegfried Nijssen arxiv

The generation of data is a common approach to improve the performance of machine learning tasks, among which is the training of models for classification. In this paper, we present TAGAL, a collection of methods able to generate synthetic tabular data using an agentic workflow. The methods leverage Large Language Models (LLMs) for an automatic and iterative process that uses feedback to improve the generated data without any further LLM training. The use of LLMs also allows for the addition of external knowledge in the generation process. We evaluate TAGAL across diverse datasets and different aspects of quality for the generated data. We look at the utility of downstream ML models, both by training classifiers on synthetic data only and by combining real and synthetic data. Moreover, we compare the similarities between the real and the generated data. We show that TAGAL is able to perform on par with state-of-the-art approaches that require LLM training and generally outperforms other training-free approaches. These findings highlight the potential of agentic workflow and open new directions for LLM-based data generation methods.

📄 PDF Abstract BibTeX arXiv:2509.04152

Code (0)

등록된 구현이 없습니다.

Tasks

Tabular Data Generation

Similar Papers 제목 키워드 기반

LightAutoDS-Tab: Multi-AutoML Agentic System for Tabular Data

2025-07-17 · Aleksey Lapin, Igor Hromov, Stanislav Chumakov, Mile Mitrovic 외 arxiv

AutoML has advanced in handling complex tasks using the integration of LLMs, yet its efficiency remains limited by dependence on specific underlying tools. In this paper, we introduce LightAutoDS-Tab, a multi-AutoML agen…

Code Generation

The UD-NewsCrawl Treebank: Reflections and Challenges from a Large-scale Tagalog Syntactic Annotation Project

2025-05-26 · Angelina A. Aquino, Lester James V. Miranda, Elsie Marie T. Or

This paper presents UD-NewsCrawl, the largest Tagalog treebank to date, containing 15.6k trees manually annotated according to the Universal Dependencies framework. We detail our treebank development process, including d…

Benchmarking zero-shot and few-shot approaches for tokenization, tagging, and dependency parsing of Tagalog text

2022-08-03 · Angelina Aquino, Franz de Leon

The grammatical analysis of texts in any written language typically involves a number of basic processing tasks, such as tokenization, morphological tagging, and dependency parsing. State-of-the-art systems can achieve h…

BenchmarkingData AugmentationDependency ParsingMorphological Tagging+1

Developing a Named Entity Recognition Dataset for Tagalog

2023-11-13 · Lester James V. Miranda

We present the development of a Named Entity Recognition (NER) dataset for Tagalog. This corpus helps fill the resource gap present in Philippine languages today, where NER resources are scarce. The texts were obtained f…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

TweetTaglish: A Dataset for Investigating Tagalog-English Code-Switching

2022-06-01 · LREC 2022 6 · Megan Herrera, Ankit Aich, Natalie Parde

Deploying recent natural language processing innovations to low-resource settings allows for state-of-the-art research findings and applications to be accessed across cultural and linguistic borders. One low-resource set…

Cultural Vocal Bursts Intensity Prediction