paper-with-me

홈 › Papers

Winning Gold at IMO 2025 with a Model-Agnostic Verification-and-Refinement Pipeline

2025-07-21 · Yichen Huang, Lin F. Yang arxiv

The International Mathematical Olympiad (IMO) is widely regarded as the world championship of high-school mathematics. IMO problems are renowned for their difficulty and novelty, demanding deep insight, creativity, and rigor. Although large language models perform well on many mathematical benchmarks, they often struggle with Olympiad-level problems. Using carefully designed prompts, we construct a model-agnostic, verification-and-refinement pipeline. We demonstrate its effectiveness on the recent IMO 2025, avoiding data contamination for models released before the competition. Equipped with any of the three leading models -- Gemini 2.5 Pro, Grok-4, or GPT-5 -- our pipeline correctly solved 5 out of the 6 problems ($\approx$85.7% accuracy). This is in sharp contrast to their baseline accuracies: 31.6% (Gemini 2.5 Pro), 21.4% (Grok-4), and 38.1% (GPT-5), obtained by selecting the best of 32 candidate solutions. The substantial improvement underscores that the path to advanced AI reasoning requires not only developing more powerful base models but also designing effective methodologies to harness their full potential for complex tasks.

📄 PDF Abstract BibTeX arXiv:2507.15855

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Decompose-and-Formalise: Recursively Verifiable Natural Language Inference

2026-01-27 · Xin Quan, Marco Valentino, Louise A. Dennis, André Freitas arxiv

Recent work has shown that integrating large language models (LLMs) with theorem provers (TPs) in neuro-symbolic pipelines helps with entailment verification and proof-guided refinement of explanations for natural langua…

Natural Language Inference

PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System

2025-09-29 · Fangchen Yu, Junchi Yao, Ziyi Wang, Haiyuan Wan 외 arxiv

Physics is central to understanding and shaping the real world, and the ability to solve physics problems is a key indicator of real-world physical intelligence. Physics Olympiads, renowned as the crown of competitive ph…

An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics

2026-09-09 · Ivan Moshkov, Stephen Ge, George Armstrong, Wei Du 외 hf

We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised…

Reinforcement Learning

PAC Verification of Statistical Algorithms

2022-11-28 · Saachi Mutreja, Jonathan Shafer

Goldwasser et al. (2021) recently proposed the setting of PAC verification, where a hypothesis (machine learning model) that purportedly satisfies the agnostic PAC learning objective is verified using an interactive proo…

PAC learning

SpecAlign: A Semantic Alignment Framework for SystemVerilog Assertion Generation

2026-05-24 · Jaime Rafael Imperial, Hao Zheng arxiv

Existing Large Language Model (LLM) approaches to SystemVerilog Assertion (SVA) generation primarily focus on syntactic validity and formal verification outcomes, while semantic alignment between generated assertions and…