paper-with-me

홈 › Papers

CodeS: Towards Building Open-source Language Models for Text-to-SQL

2024-02-26 · Haoyang Li, Jing Zhang, Hanbing Liu, Ju Fan, Xiaokang Zhang, Jun Zhu, Renjie Wei, Hongyan Pan, Cuiping Li, Hong Chen

Language models have shown promising performance on the task of translating natural language questions into SQL queries (Text-to-SQL). However, most of the state-of-the-art (SOTA) approaches rely on powerful yet closed-source large language models (LLMs), such as ChatGPT and GPT-4, which may have the limitations of unclear model architectures, data privacy risks, and expensive inference overheads. To address the limitations, we introduce CodeS, a series of pre-trained language models with parameters ranging from 1B to 15B, specifically designed for the text-to-SQL task. CodeS is a fully open-source language model, which achieves superior accuracy with much smaller parameter sizes. This paper studies the research challenges in building CodeS. To enhance the SQL generation abilities of CodeS, we adopt an incremental pre-training approach using a specifically curated SQL-centric corpus. Based on this, we address the challenges of schema linking and rapid domain adaptation through strategic prompt construction and a bi-directional data augmentation technique. We conduct comprehensive evaluations on multiple datasets, including the widely used Spider benchmark, the newly released BIRD benchmark, robustness-diagnostic benchmarks such as Spider-DK, Spider-Syn, Spider-Realistic, and Dr.Spider, as well as two real-world datasets created for financial and academic applications. The experimental results show that our CodeS achieves new SOTA accuracy and robustness on nearly all challenging text-to-SQL benchmarks.

📄 PDF Abstract BibTeX arXiv:2402.16347

Code (1)

ruckbreasoning/codes 공식 구현 pytorch

Tasks

Data AugmentationDiagnosticDomain AdaptationLanguage ModellingText to SQLText-To-SQL

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Multi-Head Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Adam 설명 없음

Similar Papers 제목 키워드 기반

Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages

2023-08-23 · Jinyi Hu, Yuan YAO, Chongyi Wang, Shan Wang 외

Recently there has been a significant surge in multimodal learning in terms of both image-to-text and text-to-image generation. However, the success is typically limited to English, leaving other languages largely behind…

Image GenerationImage to textLanguage ModelingLanguage Modelling+3

Creative Agents: Empowering Agents with Imagination for Creative Tasks

2023-12-05 · Chi Zhang, Penglin Cai, Yuhui Fu, Haoqi Yuan 외

We study building embodied agents for open-ended creative tasks. While existing methods build instruction-following agents that can perform diverse open-ended tasks, none of them demonstrates creativity -- the ability to…

Instruction FollowingLanguage ModellingLarge Language ModelMinecraft

PMC-LLaMA: Towards Building Open-source Language Models for Medicine

2023-04-27 · Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya zhang 외

Recently, Large Language Models (LLMs) have showcased remarkable capabilities in natural language understanding. While demonstrating proficiency in everyday conversations and question-answering situations, these models f…

Language ModelingLanguage ModellingMedical Question AnsweringNatural Language Understanding+1

CoDesc: A Large Code-Description Parallel Dataset

2021-05-29 · Masum Hasan, Tanveer Muttaqueen, Abdullah Al Ishtiaq, Kazi Sajeed Mehrab 외

Translation between natural language and source code can help software development by enabling developers to comprehend, ideate, search, and write computer programs in natural language. Despite growing interest from the …

Code SearchCode SummarizationSource Code SummarizationTranslation

Fine-Tuning Large Language Models and Evaluating Retrieval Methods for Improved Question Answering on Building Codes

2025-05-07 · Mohammad Aqib, Mohd Hamza, Qipei Mei, Ying Hei Chui

Building codes are regulations that establish standards for the design, construction, and safety of buildings to ensure structural integrity, fire protection, and accessibility. They are often extensive, complex, and sub…

Language ModelingLanguage ModellingNavigateQuestion Answering+3