paper-with-me

Papers

Evaluating the Text-to-SQL Capabilities of Large Language Models

2021-11-16 · ACL ARR November 2021 11 · Anonymous

We perform an empirical evaluation of Text-to-SQL capabilities of the Codex language model. We find that, without any finetuning, Codex is a strong baseline on the Spider benchmark; we also analyze the failure modes of Codex in this setting. Furthermore, we demonstrate on the GeoQuery and Scholar benchmarks that a small number of in-domain examples provided in the prompt enables Codex to perform better than state-of-the-art models finetuned on such few-shot examples.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingText to SQLText-To-SQL

Similar Papers 제목 키워드 기반

MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

2024-08-01 · Weihao Yu, Zhengyuan Yang, Lingfeng Ren, Linjie Li 외

MM-Vet, with open-ended vision-language questions targeting at evaluating integrated capabilities, has become one of the most popular benchmarks for large multimodal model evaluation. MM-Vet assesses six core vision-lang…

MathMM-VetMM-Vet v2Optical Character Recognition (OCR)+2

Evaluating the Text-to-SQL Capabilities of Large Language Models

2022-03-15 · Nitarshan Rajkumar, Raymond Li, Dzmitry Bahdanau

We perform an empirical evaluation of Text-to-SQL capabilities of the Codex language model. We find that, without any finetuning, Codex is a strong baseline on the Spider benchmark; we also analyze the failure modes of C…

Language ModelingLanguage ModellingText to SQLText-To-SQL

MMR: Evaluating Reading Ability of Large Multimodal Models

2024-08-26 · Jian Chen, Ruiyi Zhang, Yufan Zhou, Ryan Rossi 외

Large multimodal models (LMMs) have demonstrated impressive capabilities in understanding various types of image, including text-rich images. Most existing text-rich image benchmarks are simple extraction-based question …

Font RecognitionMMR totalOptical Character Recognition (OCR)Question Answering+2

CMMLU: Measuring massive multitask language understanding in Chinese

2023-06-15 · Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang 외

As the capabilities of large language models (LLMs) continue to advance, evaluating their performance becomes increasingly crucial and challenging. This paper aims to bridge this gap by introducing CMMLU, a comprehensive…

Large Language Model

CLongEval: A Chinese Benchmark for Evaluating Long-Context Large Language Models

2024-03-06 · Zexuan Qiu, Jingjing Li, Shijue Huang, Xiaoqi Jiao 외

Developing Large Language Models (LLMs) with robust long-context capabilities has been the recent research focus, resulting in the emergence of long-context LLMs proficient in Chinese. However, the evaluation of these mo…