paper-with-me

홈 › Papers

SpotIt: Evaluating Text-to-SQL Evaluation with Formal Verification

2025-10-30 · Rocky Klopfenstein, Yang He, Andrew Tremante, Yuepeng Wang, Nina Narodytska, Haoze Wu arxiv

Community-driven Text-to-SQL evaluation platforms play a pivotal role in tracking the state of the art of Text-to-SQL performance. The reliability of the evaluation process is critical for driving progress in the field. Current evaluation methods are largely test-based, which involves comparing the execution results of a generated SQL query and a human-labeled ground-truth on a static test database. Such an evaluation is optimistic, as two queries can coincidentally produce the same output on the test database while actually being different. In this work, we propose a new alternative evaluation pipeline, called SpotIt, where a formal bounded equivalence verification engine actively searches for a database that differentiates the generated and ground-truth SQL queries. We develop techniques to extend existing verifiers to support a richer SQL subset relevant to Text-to-SQL. A performance evaluation of ten Text-to-SQL methods on the high-profile BIRD dataset suggests that test-based methods can often overlook differences between the generated query and the ground-truth. Further analysis of the verification results reveals a more complex picture of the current Text-to-SQL evaluation.

📄 PDF Abstract BibTeX arXiv:2510.26840

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SpotIt+: Verification-based Text-to-SQL Evaluation with Database Constraints

2026-03-04 · Andrew Tremante, Yang He, Rocky Klopfenstein, Yuepeng Wang 외 arxiv

We present SpotIt+, an open-source tool for evaluating Text-to-SQL systems via bounded equivalence verification. Given a generated SQL query and the ground truth, SpotIt+ actively searches for database instances that dif…

FormalAlign: Automated Alignment Evaluation for Autoformalization

2024-10-14 · Jianqiao Lu, Yingjia Wan, Yinya Huang, Jing Xiong 외

Autoformalization aims to convert informal mathematical proofs into machine-verifiable formats, bridging the gap between natural and formal languages. However, ensuring semantic alignment between the informal and formali…

Mathematical Proofsvalid

Evaluating LLM-driven User-Intent Formalization for Verification-Aware Languages

2024-06-14 · Shuvendu K. Lahiri

Verification-aware programming languages such as Dafny and F* provide means to formally specify and prove properties of a program. Although the problem of checking an implementation against a specification can be defined…

Code Generationmbpp

Towards a Framework for Evaluating Explanations in Automated Fact Verification

2024-03-29 · Neema Kotonya, Francesca Toni

As deep neural models in NLP become more complex, and as a consequence opaque, the necessity to interpret them becomes greater. A burgeoning interest has emerged in rationalizing explanations to provide short and coheren…

Fact VerificationPosition

Evaluating the Safety of Deep Reinforcement Learning Models using Semi-Formal Verification

2020-10-19 · Davide Corsi, Enrico Marchesini, Alessandro Farinelli

Groundbreaking successes have been achieved by Deep Reinforcement Learning (DRL) in solving practical decision-making problems. Robotics, in particular, can involve high-cost hardware and human interactions. Hence, scrup…

Decision MakingDeep Reinforcement Learningreinforcement-learningReinforcement Learning (RL)