paper-with-me

Automated Theorem Proving 벤치마크

Automated Theorem Proving on miniF2F-test

60개 결과 · ⬇ CSV · JSON

cumulative

1.6 21.39 41.17 60.96 80.74 2021-02 2026-09 PACT (reproduced by Thor) — 24.6 (2021-02-11) PACT (reproduced by Thor) — 24.6 (2021-02-11) Lean GPT-f — 29.2 (2021-08-31) Lean tidy — 18.0 (2021-08-31) Metamath GPT-f — 1.6 (2021-08-31) Lean GPT-f — 29.2 (2021-08-31) Lean tidy — 18.0 (2021-08-31) Metamath GPT-f — 1.6 (2021-08-31) Lean Expert Iteration — 36.6 (2022-02-03) Lean Expert Iteration — 36.6 (2022-02-03) Thor — 29.9 (2022-05-22) Sledgehammer — 10.4 (2022-05-22) Thor — 29.9 (2022-05-22) Sledgehammer — 10.4 (2022-05-22) Evariste — 41.0 (2022-05-23) Evariste-7d — 40.6 (2022-05-23) Evariste-1d — 38.9 (2022-05-23) GPT-f — 36.6 (2022-05-23) Evariste — 41.0 (2022-05-23) Evariste-7d — 40.6 (2022-05-23) Evariste-1d — 38.9 (2022-05-23) GPT-f — 36.6 (2022-05-23) DSP (540B Minerva informal) — 38.9 (2022-10-21) Sledgehammer + heuristics — 20.9 (2022-10-21) DSP (540B Minerva informal) — 38.9 (2022-10-21) Sledgehammer + heuristics — 20.9 (2022-10-21) Decomposing the Enigma — 45.5 (2023-05-25) Decomposing the Enigma — 45.5 (2023-05-25) Lyra + GPT-4 — 47.1 (2023-09-27) Lyra + GPT-4 — 47.1 (2023-09-27) LEGO-Prover ChatGPT — 47.1 (2023-10-01) LEGO-Prover ChatGPT — 47.1 (2023-10-01) COPRA + GPT-4-turbo — 30.7 (2023-10-06) COPRA + GPT-4 — 23.3 (2023-10-06) COPRA + GPT-3.5 — 11.9 (2023-10-06) COPRA + GPT-4-turbo — 30.7 (2023-10-06) COPRA + GPT-4 — 23.3 (2023-10-06) COPRA + GPT-3.5 — 11.9 (2023-10-06) LLEMMA-7b — 26.2 (2023-10-16) LLEMMA-34b — 25.8 (2023-10-16) LLEMMA-7b — 26.2 (2023-10-16) LLEMMA-34b — 25.8 (2023-10-16) MMOS-DeepSeekMath-7B — 28.3 (2024-02-23) MMOS-DeepSeekMath-7B — 28.3 (2024-02-23) DeepSeek-Prover — 52.0 (2024-05-23) DeepSeek-Prover — 52.0 (2024-05-23) DeepSeek-Prover-V1.5 — 63.5 (2024-08-15) DeepSeek-Prover-V1.5 — 63.5 (2024-08-15) Subgoal-XL — 56.1 (2024-08-20) Subgoal-XL — 56.1 (2024-08-20) ProofAug — 66.0 (2025-01-30) ProofAug — 66.0 (2025-01-30) Kimina-Prover-Preview — 80.74 (2025-04-15) Kimina-Prover-Preview — 80.74 (2025-04-15) PACT (reproduced by Thor) — 24.6 (2021-02-11) Lean GPT-f — 29.2 (2021-08-31) Lean Expert Iteration — 36.6 (2022-02-03) Evariste — 41.0 (2022-05-23) Decomposing the Enigma — 45.5 (2023-05-25) Lyra + GPT-4 — 47.1 (2023-09-27) DeepSeek-Prover — 52.0 (2024-05-23) DeepSeek-Prover-V1.5 — 63.5 (2024-08-15) ProofAug — 66.0 (2025-01-30) Kimina-Prover-Preview — 80.74 (2025-04-15)
RankModel cumulativePass@1Pass@32Pass@64Pass@100ITPpass@1024pass@8192 Extra Training Data PaperCodeYear
1 Kimina-Prover-Preview 80.7452.9468.85Lean77.8780.74 Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning moonshotai/kimina-prover-preview · cmu-l3/llmlean 2025
2 ProofAug 66.036.552.5Isabelle Efficient Neural Theorem Proving via Fine-grained Proof Structure Analysis haoxiongliu/proofaug 2025
3 DeepSeek-Prover-V1.5 63.550.050.7Lean DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search deepseek-ai/deepseek-prover-v1.5 · augustepoiroux/LeanInteract 2024
4 Subgoal-XL 56.139.3Isabelle SubgoalXL: Subgoal-based Expert Learning for Theorem Proving zhaoxlpku/subgoalxl 2024
5 DeepSeek-Prover 52.030.046.3Lean DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data 2024
6 Lyra + GPT-4 47.147.1Isabelle Lyra: Orchestrating Dual Correction in Automated Theorem Proving chuanyang-zheng/lyra-theorem-prover 2023
6 LEGO-Prover ChatGPT 47.147.1Isabelle LEGO-Prover: Neural Theorem Proving with Growing Libraries wiio12/LEGO-Prover 2023
8 Decomposing the Enigma 45.545.5Isabelle Decomposing the Enigma: Subgoal-based Demonstration Learning for Formal Theorem Proving hkunlp/subgoal-theorem-prover 2023
9 Evariste 4141Lean HyperTree Proof Search for Neural Theorem Proving 2022
10 Evariste-7d 40.640.6Lean HyperTree Proof Search for Neural Theorem Proving 2022
11 Evariste-1d 38.938.9Lean HyperTree Proof Search for Neural Theorem Proving 2022
11 DSP (540B Minerva informal) 38.938.9Isabelle Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs facebookresearch/minif2f · albertqjiang/draft_sketch_prove · rah4927/lean-dojo-mew 2022
13 Lean Expert Iteration 36.629.634.536.6Lean Formal Mathematics Statement Curriculum Learning openai/lean-gym 2022
13 GPT-f 36.636.6Metamath HyperTree Proof Search for Neural Theorem Proving 2022
15 Thor + expert iteration on autoformalised theorems 35.235.2Isabelle
16 COPRA + GPT-4-turbo 30.730.7Lean An In-Context Learning Agent for Formal Theorem-Proving trishullab/copra 2023
17 Thor 29.929.9Isabelle Thor: Wielding Hammers to Integrate Language Models and Automated Theorem Provers 2022
18 Lean GPT-f 29.224.629.2Lean MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics openai/minif2f · facebookresearch/minif2f · yangky11/minif2f-lean4 · +1 2021
19 MMOS-DeepSeekMath-7B 28.328.3Lean An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning cyzhh/MMOS 2024
20 ReProver 26.526.5Lean
1–20 / 60 다음 → 페이지당 10 20 50 100