paper-with-me

Papers

Multilingual Text-to-SQL: Benchmarking the Limits of Language Models with Collaborative Language Agents

2025-09-29 · Khanh Trinh Pham, Thu Huong Nguyen, Jun Jo, Quoc Viet Hung Nguyen, Thanh Tam Nguyen arxiv

Text-to-SQL enables natural access to databases, yet most benchmarks are English-only, limiting multilingual progress. We introduce MultiSpider 2.0, extending Spider 2.0 to eight languages (English, German, French, Spanish, Portuguese, Japanese, Chinese, Vietnamese). It preserves Spider 2.0's structural difficulty while adding linguistic and dialectal variability, demanding deeper reasoning for complex SQL. On this benchmark, state-of-the-art LLMs (such as DeepSeek-R1 and OpenAI o1) reach only 4\% execution accuracy when relying on intrinsic reasoning, versus 60\% on MultiSpider 1.0. Therefore, we provide a collaboration-driven language agents baseline that iteratively refines queries, improving accuracy to 15\%. These results reveal a substantial multilingual gap and motivate methods that are robust across languages and ready for real-world enterprise deployment. Our benchmark is available at https://github.com/phkhanhtrinh23/Multilingual_Text_to_SQL.

📄 PDF Abstract BibTeX arXiv:2509.24405

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks

2026-04-21 · Peiqin Lin, Chenyang Lyu, Wenjiang Luo, Haotian Ye 외 arxiv

Large language models (LLMs) are now deployed worldwide, inspiring a surge of benchmarks that measure their multilingual and multicultural abilities. However, these benchmarks prioritize generic language understanding or…

MEGA: Multilingual Evaluation of Generative AI

2023-03-22 · Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng 외

Generative AI models have shown impressive performance on many Natural Language Processing tasks such as language understanding, reasoning, and language generation. An important question being asked by the AI community t…

Benchmarking

Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning

2025-09-26 · Antreas Ioannou, Andreas Shiamishis, Nora Hollenstein, Nezihe Merve Gürel arxiv

In an era dominated by Large Language Models (LLMs), understanding their capabilities and limitations, especially in high-stakes fields like law, is crucial. While LLMs such as Meta's LLaMA, OpenAI's ChatGPT, Google's Ge…

Adversarial RobustnessLegal Reasoning

MTG: A Benchmarking Suite for Multilingual Text Generation

2021-10-16 · ACL ARR October 2021 10 · Anonymous

We introduce MTG, a new benchmark suite for training and evaluating multilingual text generation. It is the first and largest multilingual multiway text generation benchmark with 400k human-annotated data for four tasks …

BenchmarkingQuestion GenerationQuestion-GenerationStory Generation+3

MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark

2023-10-20 · Dominik Macko, Robert Moro, Adaku Uchendu, Jason Samuel Lucas 외

There is a lack of research into capabilities of recent LLMs to generate convincing text in languages other than English and into performance of detectors of machine-generated text in multilingual settings. This is also …

Benchmarkingde-enText Detection