paper-with-me

홈 › Papers

DuSQL: A Large-Scale and Pragmatic Chinese Text-to-SQL Dataset

2020-11-01 · EMNLP 2020 11 · Lijie Wang, Ao Zhang, Kun Wu, Ke Sun, Zhenghua Li, Hua Wu, Min Zhang, Haifeng Wang

Due to the lack of labeled data, previous research on text-to-SQL parsing mainly focuses on English. Representative English datasets include ATIS, WikiSQL, Spider, etc. This paper presents DuSQL, a larges-scale and pragmatic Chinese dataset for the cross-domain text-to-SQL task, containing 200 databases, 813 tables, and 23,797 question/SQL pairs. Our new dataset has three major characteristics. First, by manually analyzing questions from several representative applications, we try to figure out the true distribution of SQL queries in real-life needs. Second, DuSQL contains a considerable proportion of SQL queries involving row or column calculations, motivated by our analysis on the SQL query distributions. Finally, we adopt an effective data construction framework via human-computer collaboration. The basic idea is automatically generating SQL queries based on the SQL grammar and constrained by the given database. This paper describes in detail the construction process and data statistics of DuSQL. Moreover, we present and compare performance of several open-source text-to-SQL parsers with minor modification to accommodate Chinese, including a simple yet effective extension to IRNet for handling calculation SQL queries.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

SQL ParsingText to SQLText-To-SQL

Similar Papers 제목 키워드 기반

CATS: A Pragmatic Chinese Answer-to-Sequence Dataset with Large Scale and High Quality

2023-06-20 · Liang Li, Ruiying Geng, Chengyang Fang, Bing Li 외

There are three problems existing in the popular data-to-text datasets. First, the large-scale datasets either contain noise or lack real application scenarios. Second, the datasets close to real applications are relativ…

Chase: A Large-Scale and Pragmatic Chinese Dataset for Cross-Database Context-Dependent Text-to-SQL

2021-08-01 · ACL 2021 5 · Jiaqi Guo, Ziliang Si, Yu Wang, Qian Liu 외

The cross-database context-dependent Text-to-SQL (XDTS) problem has attracted considerable attention in recent years due to its wide range of potential applications. However, we identify two biases in existing datasets f…

Text to SQLText-To-SQL

Data Augmentation with Hierarchical SQL-to-Question Generation for Cross-domain Text-to-SQL Parsing

2021-03-03 · EMNLP 2021 11 · Kun Wu, Lijie Wang, Zhenghua Li, Ao Zhang 외

Data augmentation has attracted a lot of research attention in the deep learning era for its ability in alleviating data sparseness. The lack of labeled data for unseen evaluation databases is exactly the major challenge…

Data AugmentationQuestion GenerationQuestion-GenerationSQL Parsing+2

The Mask of Civility: Benchmarking Chinese Mock Politeness Comprehension in Large Language Models

2026-02-03 · Yitong Zhang, Yuhan Xiang, Mingxuan Liu arxiv

From a pragmatic perspective, this study systematically evaluates the differences in performance among representative large language models (LLMs) in recognizing politeness, impoliteness, and mock politeness phenomena in…

Generating Bilingual Pragmatic Color References

2018-03-11 · NAACL 2018 6 · Will Monroe, Jennifer Hu, Andrew Jong, Christopher Potts

Contextual influences on language often exhibit substantial cross-lingual regularities; for example, we are more verbose in situations that require finer distinctions. However, these regularities are sometimes obscured b…