paper-with-me

Papers

Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis

2025-11-28 · Matej Klemen, Tjaša Arčon, Luka Terčon, Marko Robnik-Šikonja, Kaja Dobrovoljc arxiv

Empirical grammar research has become increasingly data-driven, but the systematic analysis of annotated corpora still requires substantial methodological and technical effort. We explore how agentic large language models (LLMs) can streamline this process by reasoning over annotated corpora and producing interpretable, data-grounded answers to linguistic questions. We introduce an agentic framework for corpus-grounded grammatical analysis that integrates concepts such as natural-language task interpretation, code generation, and data-driven reasoning. As a proof of concept, we apply it to Universal Dependencies (UD) corpora, testing it on multilingual grammatical tasks inspired by the World Atlas of Language Structures (WALS). The evaluation spans 13 word-order features and over 170 languages, assessing system performance across three complementary dimensions - dominant-order accuracy, order-coverage completeness, and distributional fidelity - which reflect how well the system generalizes, identifies, and quantifies word-order variations. The results demonstrate the feasibility of combining LLM reasoning with structured linguistic data, offering a first step toward interpretable, scalable automation of corpus-based grammatical inquiry.

📄 PDF Abstract BibTeX arXiv:2512.00214

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

From RAG to Agentic RAG for Faithful Islamic Question Answering

2026-01-12 · Gagan Bhatia, Hamdy Mubarak, Mustafa Jarrar, George Mikros 외 arxiv

Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences. Yet standard MCQ/MRC-style evaluations (MCQ: Multiple choice questio…

Machine Reading ComprehensionQuestion Answering

BiST: A Gold Standard Bangla-English Bilingual Corpus for Sentence Structure and Tense Classification with Inter-Annotator Agreement

2026-04-06 · Abdullah Al Shafi, Swapnil Kundu Argha, M. A. Moyeen, Abdul Muntakim 외 arxiv

High-quality bilingual resources remain a critical bottleneck for advancing multilingual NLP in low-resource settings, particularly for Bangla. To mitigate this gap, we introduce BiST, a rigorously curated Bangla-English…

Representation LearningLanguage IdentificationText Generation

MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation

2026-04-20 · Mehul Agarwal, Aditya Aggarwal, Arnav Goel, Medha Hira 외 arxiv

While multilingual large language models (LLMs) perform well on high-level tasks like translation and question answering, their ability to handle grammatical gender and morphological agreement remains underexplored. In m…

Question Answering

"Be My Cheese?": Cultural Nuance Benchmarking for Machine Translation in Multilingual LLMs

2026-02-04 · Madison Van Doren, Casey Ford, Jennifer Barajas, Riley VanMeter 외 arxiv

We present a large-scale human evaluation benchmark for assessing cultural localisation in machine translation produced by state-of-the-art multilingual large language models (LLMs). Existing MT benchmarks emphasise toke…

Machine Translation

MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights

2026-06-05 · Yilun Liu, Miao Zhang, Shimin Tao, Minggui He 외 arxiv

Multilingual and multicultural benchmarks now cover dozens of languages and model families, but the resulting score landscapes remain metric-rich and insight-poor, necessitating fine-grained multilingual post-evaluation …