paper-with-me

홈 › Papers

Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL

2025-04-21 · Simone Papicchio, Simone Rossi, Luca Cagliero, Paolo Papotti

Large Language Models (LLMs) have shown impressive capabilities in transforming natural language questions about relational databases into SQL queries. Despite recent improvements, small LLMs struggle to handle questions involving multiple tables and complex SQL patterns under a Zero-Shot Learning (ZSL) setting. Supervised Fine-Tuning (SFT) partially compensates for the knowledge deficits in pretrained models but falls short while dealing with queries involving multi-hop reasoning. To bridge this gap, different LLM training strategies to reinforce reasoning capabilities have been proposed, ranging from leveraging a thinking process within ZSL, including reasoning traces in SFT, or adopt Reinforcement Learning (RL) strategies. However, the influence of reasoning on Text2SQL performance is still largely unexplored. This paper investigates to what extent LLM reasoning capabilities influence their Text2SQL performance on four benchmark datasets. To this end, it considers the following LLM settings: (1) ZSL, including general-purpose reasoning or not; (2) SFT, with and without task-specific reasoning traces; (3) RL, exploring the use of different rewarding functions, both the established EXecution accuracy (EX) and a mix with fine-grained ones that also account the precision, recall, and cardinality of partially correct answers; (4) SFT+RL, i.e, a two-stage approach that combines SFT and RL. The results show that general-purpose reasoning under ZSL proves to be ineffective in tackling complex Text2SQL cases. Small LLMs benefit from SFT with reasoning much more than larger ones. RL is generally beneficial across all tested models and datasets. The use of the fine-grained metrics turns out to be the most effective RL strategy. Thanks to RL and the novel text2SQL rewards, the 7B Qwen-Coder-2.5 model performs on par with 400+ Billion ones (including gpt-4o) on the Bird dataset.

📄 PDF Abstract BibTeX arXiv:2504.15077

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning (RL)Zero-Shot Learning

Methods 이 논문이 사용한 방법론

SFT Shrink and Fine-Tune, or SFT, is a type of distillation that avoids explicit distillation by copying parameters to a student student model and then fine-tuning.…
ADOPT Please enter a description about the method here

Similar Papers 제목 키워드 기반

Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning

2025-08-11 · Shu Wu, Chenxing Li, Wenfu Wang, Hao Zhang 외 arxiv

Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning with rule-ba…

Reinforcement LearningQuestion Answering

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs

2025-07-03 · Purbesh Mitra, Sennur Ulukus arxiv

Recent advancements in the reasoning capabilities of large language models (LLMs) show that employing group relative policy optimization (GRPO) algorithm for reinforcement learning (RL) training allows the models to use …

Reinforcement Learning

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking

2025-05-25 · Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari

Recent advances in large language models (LLMs) demonstrate their impressive reasoning capabilities. However, the reasoning confined to internal parametric space limits LLMs' access to real-time information and understan…

Mathematical ReasoningMulti-hop Question AnsweringQuestion Answeringtext-based games

TKG-Thinker: Towards Dynamic Reasoning over Temporal Knowledge Graphs via Agentic Reinforcement Learning

2026-02-05 · Zihao Jiang, Miao Peng, Zhenyan Shan, Wenjie Xu 외 arxiv

Temporal knowledge graph question answering (TKGQA) aims to answer time-sensitive questions by leveraging temporal knowledge bases. While Large Language Models (LLMs) demonstrate significant potential in TKGQA, current p…

Graph Question AnsweringReinforcement LearningKnowledge Graphs

VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning

2026-03-15 · Chaoyang Wang, Wenrui Bao, Sicheng Gao, Bingxin Xu 외 arxiv

Vision-Language-Action (VLA) models have shown promising capabilities for embodied intelligence, but most existing approaches rely on text-based chain-of-thought reasoning where visual inputs are treated as static contex…

Reinforcement Learning