paper-with-me

Papers

GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs

2025-08-19 · Adrian-Marius Dumitran, Alexandra-Mihaela Danila, Angela-Liliana Dumitran arxiv

LLMs (Large language models) have revolutionized NLP (Natural Language Processing), yet their pedagogical value for low-resource languages remains unclear. We present GRILE (Grammar Romanian Inference and Language Explanations) , the first open benchmark of 1,151 multiple-choice questions harvested from Romanian high-stakes exams (National Evaluation, Baccalaureate, university admissions). GRILE enables us to probe two complementary abilities of seven state-of-the-art multilingual and Romanian-specific LLMs: (i) selecting the correct answer, and (ii) producing linguistically accurate explanations. While Gemini 2.5 Pro reaches 83% accuracy, most open-weight models stay below 65%, and 48% of their explanations contain factual or pedagogical flaws according to expert review. A detailed error analysis pinpoints systematic weaknesses in morphology and in applying the latest DOOM3 orthographic norms. All data, code and a public web demo are released to catalyze future research. Our findings expose open challenges for trustworthy educational NLP in low-resource settings and establish GRILE as a new test-bed for controllable explanation generation and evaluation.

📄 PDF Abstract BibTeX arXiv:2508.14279

Code (0)

등록된 구현이 없습니다.

Tasks

Explanation Generation

Similar Papers 제목 키워드 기반

RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams

2025-07-25 · Andrei Vlad Man, Răzvan-Alexandru Smădu, Cristian-George Craciun, Dumitru-Clementin Cercel 외 arxiv

The intersection of AI and legal systems presents a growing need for tools that support legal education, particularly in under-resourced languages such as Romanian. In this work, we aim to evaluate the capabilities of La…

Information RetrievalQuestion Answering

Romanian TimeBank: An Annotated Parallel Corpus for Temporal Information

2012-05-01 · LREC 2012 5 · Corina For{\u{a}}scu, Dan Tufi{\c{s}}

The paper describes the main steps for the construction, annotation and validation of the Romanian version of the TimeBank corpus. Starting from the English TimeBank corpus ― the reference annotated corpus in the tempo…

Information RetrievalMachine TranslationQuestion AnsweringTAG

CRANE: Reasoning with constrained LLM generation

2025-02-13 · Debangshu Banerjee, Tarun Suresh, Shubham Ugare, Sasa Misailovic 외

Code generation, symbolic math reasoning, and other tasks require LLMs to produce outputs that are both syntactically and semantically correct. Constrained LLM generation is a promising direction to enforce adherence to …

Code GenerationMathvalid

RoMath: A Mathematical Reasoning Benchmark in Romanian

2024-09-17 · Adrian Cosma, Ana-Maria Bucur, Emilian Radoi

Mathematics has long been conveyed through natural language, primarily for human understanding. With the rise of mechanized mathematics and proof assistants, there is a growing need to understand informal mathematical te…

Mathematical Reasoning

TF3-RO-50M: Training Compact Romanian Language Models from Scratch on Synthetic Moral Microfiction

2026-01-15 · Mihai Dan Nadas, Laura Diosan, Andreea Tomescu, Andrei Piscoran arxiv

Recent advances in synthetic data generation have shown that compact language models can be trained effectively when the underlying corpus is structurally controlled and linguistically coherent. However, for morphologica…

Synthetic Data GenerationKnowledge Distillation