paper-with-me

Papers

RoMath: A Mathematical Reasoning Benchmark in Romanian

2024-09-17 · Adrian Cosma, Ana-Maria Bucur, Emilian Radoi

Mathematics has long been conveyed through natural language, primarily for human understanding. With the rise of mechanized mathematics and proof assistants, there is a growing need to understand informal mathematical text, yet most existing benchmarks focus solely on English, overlooking other languages. This paper introduces RoMath, a Romanian mathematical reasoning benchmark suite comprising three subsets: Baccalaureate, Competitions and Synthetic, which cover a range of mathematical domains and difficulty levels, aiming to improve non-English language models and promote multilingual AI development. By focusing on Romanian, a low-resource language with unique linguistic features, RoMath addresses the limitations of Anglo-centric models and emphasizes the need for dedicated resources beyond simple automatic translation. We benchmark several open-weight language models, highlighting the importance of creating resources for underrepresented languages. Code and datasets are be made available.

📄 PDF Abstract BibTeX arXiv:2409.11074

Code (1)

cosmaadrian/romath 공식 구현

Tasks

Mathematical Reasoning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)

2026-03-28 · Luca-Ncolae Cuclea, Sabin-Codrut Badea, Adrian-Marius Dumitran arxiv

AI in Education research increasingly relies on authentic, curriculum-grounded assessment data, yet large, well-structured exam corpora remain scarce for many languages and educational systems. We introduce RoMathExam, a…

RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams

2025-07-25 · Andrei Vlad Man, Răzvan-Alexandru Smădu, Cristian-George Craciun, Dumitru-Clementin Cercel 외 arxiv

The intersection of AI and legal systems presents a growing need for tools that support legal education, particularly in under-resourced languages such as Romanian. In this work, we aim to evaluate the capabilities of La…

Information RetrievalQuestion Answering

GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs

2025-08-19 · Adrian-Marius Dumitran, Alexandra-Mihaela Danila, Angela-Liliana Dumitran arxiv

LLMs (Large language models) have revolutionized NLP (Natural Language Processing), yet their pedagogical value for low-resource languages remains unclear. We present GRILE (Grammar Romanian Inference and Language Explan…

Explanation Generation

A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian

2025-08-22 · Ana-Cristina Rogoz, Radu Tudor Ionescu, Alexandra-Valentina Anghel, Ionut-Lucian Antone-Iordache 외 arxiv

We introduce MedQARo, the first large-scale medical QA benchmark in Romanian, alongside a comprehensive evaluation of state-of-the-art large language models (LLMs). We construct a high-quality and large-scale dataset com…

Keyword ExtractionQuestion Answering

Exploring Large Language Models for Translating Romanian Computational Problems into English

2025-01-09 · Adrian Marius Dumitran, Adrian-Catalin Badea, Stefan-Gabriel Muscalu, Angela-Liliana Dumitran 외

Recent studies have suggested that large language models (LLMs) underperform on mathematical and computer science tasks when these problems are translated from Romanian into English, compared to their original Romanian f…

Translation