paper-with-me

홈 › Papers

Solving a Research Problem in Mathematical Statistics with AI Assistance

2025-11-24 · Edgar Dobriban arxiv

Over the last few months, AI models including large language models have improved greatly. There are now several documented examples where they have helped professional mathematical scientists prove new results, sometimes even helping resolve known open problems. In this short note, we add another example to the list, by documenting how we were able to solve a previously unsolved research problem in robust mathematical statistics with crucial help from GPT-5. Our problem concerns robust density estimation, where the observations are perturbed by Wasserstein-bounded contaminations. In a previous preprint (Chao and Dobriban, 2023, arxiv:2308.01853v2), we have obtained upper and lower bounds on the minimax optimal estimation error; which were, however, not sharp. Starting in October 2025, making significant use of GPT-5 Pro, we were able to derive the minimax optimal error rate (reported in version 3 of the above arxiv preprint). GPT-5 provided crucial help along the way, including by suggesting calculations that we did not think of, and techniques that were not familiar to us, such as the dynamic Benamou-Brenier formulation, for key steps in the analysis. Working with GPT-5 took a few weeks of effort, and we estimate that it could have taken several months to get the same results otherwise. At the same time, there are still areas where working with GPT-5 was challenging: it sometimes provided incorrect references, and glossed over details that sometimes took days of work to fill in. We outline our workflow and steps taken to mitigate issues. Overall, our work can serve as additional documentation for a new age of human-AI collaborative work in mathematical science.

📄 PDF Abstract BibTeX arXiv:2511.18828

Code (0)

등록된 구현이 없습니다.

Tasks

Density Estimation

Similar Papers 제목 키워드 기반

MACM: Utilizing a Multi-Agent System for Condition Mining in Solving Complex Mathematical Problems

2024-04-06 · Bin Lei, Yi Zhang, Shan Zuo, Ali Payani 외

Recent advancements in large language models, such as GPT-4, have demonstrated remarkable capabilities in processing standard queries. Despite these advancements, their performance substantially declines in \textbf{advan…

Logical ReasoningMathMath Word Problem Solving

Mask-Proof: An LLM-based Automated Data Curation Pipeline on Mathematical Proofs

2026-06-13 · Jierui Zhang, Siyuan Tan, Xinhang Li, Longzhuangzhi Lin 외 arxiv

Large language models (LLMs) are increasingly capable of mathematical problem solving and can even assist with research-level proofs, yet we still lack a scalable and reproducible way to measure step-level reasoning in l…

Mathematical Reasoning

HintMR: Eliciting Stronger Mathematical Reasoning in Small Language Models

2026-04-14 · Jawad Hossain, Xiangyu Guo, Jiawei Zhou, Chong Liu arxiv

Small language models (SLMs) often struggle with complex mathematical reasoning due to limited capacity to maintain long chains of intermediate steps and to recover from early errors. We address this challenge by introdu…

Mathematical Reasoning

CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images

2025-10-13 · Chengqi Duan, Kaiyue Sun, Rongyao Fang, Manyuan Zhang 외 arxiv

Recent advances in Large Language Models (LLMs) and Vision Language Models (VLMs) have shown significant progress in mathematical reasoning, yet they still face a critical bottleneck with problems requiring visual assist…

Mathematical ReasoningVisual Reasoning

Measuring Mathematical Problem Solving With the MATH Dataset

2021-03-05 · Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora 외

Many intellectual endeavors require mathematical problem solving, but this skill remains beyond the capabilities of computers. To measure this ability in machine learning models, we introduce MATH, a new dataset of 12,50…

MathMathematical Problem-SolvingMathematical ReasoningMath Word Problem Solving+1