paper-with-me

홈 › Papers

Semantic Evaluation for Text-to-SQL with Distilled Test Suites

2020-10-06 · EMNLP 2020 11 · Ruiqi Zhong, Tao Yu, Dan Klein

We propose test suite accuracy to approximate semantic accuracy for Text-to-SQL models. Our method distills a small test suite of databases that achieves high code coverage for the gold query from a large number of randomly generated databases. At evaluation time, it computes the denotation accuracy of the predicted queries on the distilled test suite, hence calculating a tight upper-bound for semantic accuracy efficiently. We use our proposed method to evaluate 21 models submitted to the Spider leader board and manually verify that our method is always correct on 100 examples. In contrast, the current Spider metric leads to a 2.5% false negative rate on average and 8.1% in the worst case, indicating that test suite accuracy is needed. Our implementation, along with distilled test suites for eleven Text-to-SQL datasets, is publicly available.

📄 PDF Abstract BibTeX arXiv:2010.02840

Code (3)

ruiqi-zhong/TestSuiteEval 공식 구현
taoyds/spider tf
taoyds/test-suite-sql-eval

Tasks

Text to SQLText-To-SQL

Similar Papers 제목 키워드 기반

Semantic Evaluation for Text-to-SQL with Distilled Test Suite

2020-07-02 · Ruiqi Zhong, Tao Yu, Dan Klein

We propose test suite accuracy to approximate semantic accuracy for Text-to-SQL models, where a predicted query is semantically correct if its denotation is the same as the gold for every possible database. Our method di…

Semantic ParsingText to SQLText-To-SQL

Test Suites Task: Evaluation of Gender Fairness in MT with MuST-SHE and INES

2023-10-30 · Beatrice Savoldi, Marco Gaido, Matteo Negri, Luisa Bentivogli

As part of the WMT-2023 "Test suites" shared task, in this paper we summarize the results of two test suites evaluations: MuST-SHE-WMT23 and INES. By focusing on the en-de and de-en language pairs, we rely on these newly…

de-enFairness

RBT4DNN: Requirements-based Testing of Neural Networks

2025-04-03 · Nusrat Jahan Mozumder, Felipe Toledo, Swaroopa Dola, Matthew B. Dwyer

Deep neural network (DNN) testing is crucial for the reliability and safety of critical systems, where failures can have severe consequences. Although various techniques have been developed to create robustness test suit…

DNN Testing

Distilling Facial Knowledge With Teacher-Tasks: Semantic-Segmentation-Features For Pose-Invariant Face-Recognition

2022-09-02 · Ali Hassani, Zaid El Shair, Rafi Ud Duala Refat, Hafiz Malik

This paper demonstrates a novel approach to improve face-recognition pose-invariance using semantic-segmentation features. The proposed Seg-Distilled-ID network jointly learns identification and semantic-segmentation tas…

Face RecognitionRobust Face RecognitionSegmentationSemantic Segmentation

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications

2026-01-29 · Daniel Commey arxiv

Evaluating Large Language Model (LLM) applications differs from conventional software testing because outputs are probabilistic, semantically variable, and sensitive to prompt and model changes. This technical report pro…