paper-with-me

홈 › Papers

DOoM: Difficult Olympiads of Math

2025-09-27 · Ilya Kuleshov, Ilin Pavel, Nikolay Kompanets, Ksenia Sycheva, Aleksandr Nikolich arxiv

This paper introduces DOoM, a new open-source benchmark designed to assess the capabilities of language models in solving mathematics and physics problems in Russian. The benchmark includes problems of varying difficulty, ranging from school-level tasks to university Olympiad and entrance exam questions. In this paper we discuss the motivation behind its creation, describe dataset's structure and evaluation methodology, and present initial results from testing various models. Analysis of the results shows a correlation between model performance and the number of tokens used, and highlights differences in performance between mathematics and physics tasks.

📄 PDF Abstract BibTeX arXiv:2509.23529

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Formal Mathematics Statement Curriculum Learning

2022-02-03 · Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys 외

We explore the use of expert iteration in the context of language modeling applied to formal mathematics. We show that at same compute budget, expert iteration, by which we mean proof search interleaved with learning, dr…

Automated Theorem ProvingLanguage ModelingLanguage Modelling

Proposing and solving olympiad geometry with guided tree search

2024-12-14 · Chi Zhang, Jiajun Song, Siyu Li, Yitao Liang 외

Mathematics olympiads are prestigious competitions, with problem proposing and solving highly honored. Building artificial intelligence that proposes and solves olympiads presents an unresolved challenge in automated the…

EEFSUVA: A New Mathematical Olympiad Benchmark

2025-09-23 · Nicole N Khatibi, Daniil A. Radamovich, Michael P. Brenner arxiv

Recent breakthroughs have spurred claims that large language models (LLMs) match gold medal Olympiad to graduate level proficiency on mathematics benchmarks. In this work, we examine these claims in detail and assess the…

Mathematical Reasoning

Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads

2024-06-22 · Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Joanna Matthiesen 외

Recent years have seen a significant progress in the general-purpose problem solving abilities of large vision and language models (LVLMs), such as ChatGPT, Gemini, etc.; some of these breakthroughs even seem to enable A…

Mathematical Reasoning

Learning Latent Graph Dynamics for Visual Manipulation of Deformable Objects

2021-04-25 · Xiao Ma, David Hsu, Wee Sun Lee

Manipulating deformable objects, such as ropes and clothing, is a long-standing challenge in robotics, because of their large degrees of freedom, complex non-linear dynamics, and self-occlusion in visual perception. The …

Contrastive LearningDeformable Object ManipulationGraph Neural NetworkModel Predictive Control+2