Makadi: A Large-Scale Human-Labeled Dataset for Hindi Semantic Parsing
Parsing natural language queries into formal database calls is a very well-studied problem. Because of the rich diversity of semantic markers across the world’s languages, progress in solving this problem is irreducibly language-dependent. This has created an asymmetry in progress in NLIDB solutions, with most state-of-the-art efforts focused on the resource-rich English language, with limited progress seen for low resource languages. In this short paper, we present Makadi, a large-scale, complex, cross-lingual, cross-domain semantic parsing and text-to-SQL dataset for semantic parsing in the Hindi language. Produced by translating the recently introduced English language Spider NLIDB dataset, it consists of 9693 questions and SQL queries on 166 databases with multiple tables which cover multiple domains. This is the first large-scale dataset in the Hindi language for semantic parsing and related language understanding tasks. Our dataset is publicly available at: Link removed to preserve anonymization during peer review.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityNatural Language QueriesSemantic ParsingText to SQLText-To-SQLSimilar Papers 제목 키워드 기반
Large-scale Datasets: Faces with Partial Occlusions and Pose Variations in the Wild
Face detection methods have relied on face datasets for training. However, existing face datasets tend to be in small scales for face learning in both constrained and unconstrained environments. In this paper, we first i…
Face DetectionFace RecognitionExploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
The vocabulary size in temporal action localization (TAL) is limited by the scarcity of large-scale annotated datasets. To overcome this, recent works integrate vision-language models (VLMs), such as CLIP, for open-vocab…
Action LocalizationTemporal Action LocalizationDomainMix: Learning Generalizable Person Re-Identification Without Human Annotations
Existing person re-identification models often have low generalizability, which is mostly due to limited availability of large-scale labeled data in training. However, labeling large-scale training data is very expensive…
Domain AdaptationGeneralizable Person Re-identificationPerson Re-IdentificationUnsupervised Domain AdaptationAutomatically Labeled Data Generation for Large Scale Event Extraction
Modern models of event extraction for tasks like ACE are based on supervised learning of events from small hand-labeled data. However, hand-labeled training data is expensive to produce, in low coverage of event types, a…
Event ExtractionKnowledge Base PopulationRelation ExtractionWorld KnowledgeMATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization
Recently, large-scale datasets have vastly facilitated the development in nearly all domains of Natural Language Processing. However, there is currently no cross-task dataset in NLP, which hinders the development of mult…
ClassificationGeneral ClassificationMulti-Task LearningQuestion Answering