paper-with-me

Papers

Makadi: A Large-Scale Human-Labeled Dataset for Hindi Semantic Parsing

2022-06-01 · WILDRE (LREC) 2022 6 · Shashwat Vaibhav, Nisheeth Srivastava

Parsing natural language queries into formal database calls is a very well-studied problem. Because of the rich diversity of semantic markers across the world’s languages, progress in solving this problem is irreducibly language-dependent. This has created an asymmetry in progress in NLIDB solutions, with most state-of-the-art efforts focused on the resource-rich English language, with limited progress seen for low resource languages. In this short paper, we present Makadi, a large-scale, complex, cross-lingual, cross-domain semantic parsing and text-to-SQL dataset for semantic parsing in the Hindi language. Produced by translating the recently introduced English language Spider NLIDB dataset, it consists of 9693 questions and SQL queries on 166 databases with multiple tables which cover multiple domains. This is the first large-scale dataset in the Hindi language for semantic parsing and related language understanding tasks. Our dataset is publicly available at: Link removed to preserve anonymization during peer review.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityNatural Language QueriesSemantic ParsingText to SQLText-To-SQL

Similar Papers 제목 키워드 기반

Large-scale Datasets: Faces with Partial Occlusions and Pose Variations in the Wild

2017-06-27 · Tarik Alafif, Zeyad Hailat, Melih Aslan, Xue-wen Chen

Face detection methods have relied on face datasets for training. However, existing face datasets tend to be in small scales for face learning in both constrained and unconstrained environments. In this paper, we first i…

Face DetectionFace Recognition

Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization

2024-07-09 · Jeongseok Hyun, Su Ho Han, Hyolim Kang, Joon-Young Lee 외

The vocabulary size in temporal action localization (TAL) is limited by the scarcity of large-scale annotated datasets. To overcome this, recent works integrate vision-language models (VLMs), such as CLIP, for open-vocab…

Action LocalizationTemporal Action Localization

DomainMix: Learning Generalizable Person Re-Identification Without Human Annotations

2020-11-24 · Wenhao Wang, Shengcai Liao, Fang Zhao, Cuicui Kang 외

Existing person re-identification models often have low generalizability, which is mostly due to limited availability of large-scale labeled data in training. However, labeling large-scale training data is very expensive…

Domain AdaptationGeneralizable Person Re-identificationPerson Re-IdentificationUnsupervised Domain Adaptation

Automatically Labeled Data Generation for Large Scale Event Extraction

2017-07-01 · ACL 2017 7 · Yubo Chen, Shulin Liu, Xiang Zhang, Kang Liu 외

Modern models of event extraction for tasks like ACE are based on supervised learning of events from small hand-labeled data. However, hand-labeled training data is expensive to produce, in low coverage of event types, a…

Event ExtractionKnowledge Base PopulationRelation ExtractionWorld Knowledge

MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization

2020-04-26 · ACL 2020 6 · Canwen Xu, Jiaxin Pei, Hongtao Wu, Yiyu Liu 외

Recently, large-scale datasets have vastly facilitated the development in nearly all domains of Natural Language Processing. However, there is currently no cross-task dataset in NLP, which hinders the development of mult…

ClassificationGeneral ClassificationMulti-Task LearningQuestion Answering