Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
A significant amount of the world's knowledge is stored in relational databases. However, the ability for users to retrieve facts from a database is limited due to a lack of understanding of query languages such as SQL. We propose Seq2SQL, a deep neural network for translating natural language questions to corresponding SQL queries. Our model leverages the structure of SQL queries to significantly reduce the output space of generated queries. Moreover, we use rewards from in-the-loop query execution over the database to learn a policy to generate unordered parts of the query, which we show are less suitable for optimization via cross entropy loss. In addition, we will publish WikiSQL, a dataset of 80654 hand-annotated examples of questions and SQL queries distributed across 24241 tables from Wikipedia. This dataset is required to train our model and is an order of magnitude larger than comparable datasets. By applying policy-based reinforcement learning with a query execution environment to WikiSQL, our model Seq2SQL outperforms attentional sequence to sequence models, improving execution accuracy from 35.9% to 59.4% and logical form accuracy from 23.4% to 48.3%.
Code (15)
Tasks
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Text-To-SQLSimilar Papers 제목 키워드 기반
Learning to Generate Structured Queries from Natural Language with Indirect Supervision
Generating structured query language (SQL) from natural language is an emerging research topic. This paper presents a new learning paradigm from indirect supervision of the answers to natural language questions, instead …
reinforcement-learningReinforcement LearningReinforcement Learning (RL)SQLNet: Generating Structured Queries From Natural Language Without Reinforcement Learning
Synthesizing SQL queries from natural language is a long-standing open problem and has been attracting considerable interest recently. Toward solving the problem, the de facto approach is to employ a sequence-to-sequence…
Decoderreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1LAFA: Agentic LLM-Driven Federated Analytics over Decentralized Data Sources
Large Language Models (LLMs) have shown great promise in automating data analytics tasks by interpreting natural language queries and generating multi-operation execution plans. However, existing LLM-agent-based analytic…
Natural Language QueriesQ-NL Verifier: Leveraging Synthetic Data for Robust Knowledge Graph Question Answering
Question answering (QA) requires accurately aligning user questions with structured queries, a process often limited by the scarcity of high-quality query-natural language (Q-NL) pairs. To overcome this, we present Q-NL …
Graph Question AnsweringQuestion AnsweringTranslationNatural Language Interface for Databases Using a Dual-Encoder Model
We propose a sketch-based two-step neural model for generating structured queries (SQL) based on a user{'}s request in natural language. The sketch is obtained by using placeholders for specific entities in the SQL query…
Machine TranslationSemantic Parsing