paper-with-me

홈 › Papers

Data Augmentation with Hierarchical SQL-to-Question Generation for Cross-domain Text-to-SQL Parsing

2021-03-03 · EMNLP 2021 11 · Kun Wu, Lijie Wang, Zhenghua Li, Ao Zhang, Xinyan Xiao, Hua Wu, Min Zhang, Haifeng Wang

Data augmentation has attracted a lot of research attention in the deep learning era for its ability in alleviating data sparseness. The lack of labeled data for unseen evaluation databases is exactly the major challenge for cross-domain text-to-SQL parsing. Previous works either require human intervention to guarantee the quality of generated data, or fail to handle complex SQL queries. This paper presents a simple yet effective data augmentation framework. First, given a database, we automatically produce a large number of SQL queries based on an abstract syntax tree grammar. For better distribution matching, we require that at least 80% of SQL patterns in the training data are covered by generated queries. Second, we propose a hierarchical SQL-to-question generation model to obtain high-quality natural language questions, which is the major contribution of this work. Finally, we design a simple sampling strategy that can greatly improve training efficiency given large amounts of generated data. Experiments on three cross-domain datasets, i.e., WikiSQL and Spider in English, and DuSQL in Chinese, show that our proposed data augmentation framework can consistently improve performance over strong baselines, and the hierarchical generation component is the key for the improvement.

📄 PDF Abstract BibTeX arXiv:2103.02227

Code (1)

PaddlePaddle/Research 공식 구현 paddle

Tasks

Data AugmentationQuestion GenerationQuestion-GenerationSQL ParsingText to SQLText-To-SQL

Similar Papers 제목 키워드 기반

Tell Me How to Ask Again: Question Data Augmentation with Controllable Rewriting in Continuous Space

2020-10-04 · EMNLP 2020 11 · Dayiheng Liu, Yeyun Gong, Jie Fu, Yu Yan 외

In this paper, we propose a novel data augmentation method, referred to as Controllable Rewriting based Question Data Augmentation (CRQDA), for machine reading comprehension (MRC), question generation, and question-answe…

Data AugmentationMachine Reading ComprehensionNatural Language InferenceQNLI+5

ZusammenQA: Data Augmentation with Specialized Models for Cross-lingual Open-retrieval Question Answering System

2022-05-30 · NAACL (MIA) 2022 7 · Chia-Chien Hung, Tommaso Green, Robert Litschko, Tornike Tsereteli 외

This paper introduces our proposed system for the MIA Shared Task on Cross-lingual Open-retrieval Question Answering (COQA). In this challenging scenario, given an input question the system has to gather evidence documen…

Answer GenerationData AugmentationLanguage ModellingPassage Retrieval+2

Question Generation from Paragraphs: A Tale of Two Hierarchical Models

2019-11-08 · Vishwajeet Kumar, Raktim Chaki, Sai Teja Talluri, Ganesh Ramakrishnan 외

Automatic question generation from paragraphs is an important and challenging problem, particularly due to the long context from paragraphs. In this paper, we propose and study two hierarchical models for the task of que…

Question GenerationQuestion-GenerationSentenceVocal Bursts Valence Prediction

Balancing Cost and Effectiveness of Synthetic Data Generation Strategies for LLMs

2024-09-29 · Yung-Chieh Chan, George Pu, Apaar Shanker, Parth Suresh 외

As large language models (LLMs) are applied to more use cases, creating high quality, task-specific datasets for fine-tuning becomes a bottleneck for model improvement. Using high quality human data has been the most com…

Synthetic Data Generation

HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA

2024-02-01 · Xinyue Chen, Pengyu Gao, Jiangjiang Song, Xiaoyang Tan

Retrieval-augmented generation (RAG) has rapidly advanced the language model field, particularly in question-answering (QA) systems. By integrating external documents during the response generation phase, RAG significant…

HallucinationLanguage ModelingLanguage ModellingQuestion Answering+4