paper-with-me

홈 › Papers

Learning to Ideate for Machine Learning Engineering Agents

2026-01-24 · Yunxiang Zhang, Kang Zhou, Zhichao Xu, Kiran Ramnath, Yun Zhou, Sangmin Woo, Haibo Ding, Lin Lee Cheong arxiv

Existing machine learning engineering (MLE) agents struggle to iteratively optimize their implemented algorithms for effectiveness. To address this, we introduce MLE-Ideator, a dual-agent framework that separates ideation from implementation. In our system, an implementation agent can request strategic help from a dedicated Ideator. We show this approach is effective in two ways. First, in a training-free setup, our framework significantly outperforms implementation-only agent baselines on MLE-Bench. Second, we demonstrate that the Ideator can be trained with reinforcement learning (RL) to generate more effective ideas. With only 1K training samples from 10 MLE tasks, our RL-trained Qwen3-8B Ideator achieves an 11.5% relative improvement compared to its untrained counterpart and surpasses Claude Sonnet 3.5. These results highlights a promising path toward training strategic AI systems for scientific discovery.

📄 PDF Abstract BibTeX arXiv:2601.17596

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Agent Ideate: A Framework for Product Idea Generation from Patents Using Agentic AI

2025-07-02 · Gopichand Kanumolu, Ashok Urlana, Charaka Vinayak Kumar, Bala Mallikarjunarao Garlapati arxiv

Patents contain rich technical knowledge that can inspire innovative product ideas, yet accessing and interpreting this information remains a challenge. This work explores the use of Large Language Models (LLMs) and auto…

Imagining Design Workflows in Agentic AI Futures

2025-09-25 · Samangi Wadinambiarachchi, Jenny Waycott, Yvonne Rogers, Greg Wadley arxiv

As designers become familiar with Generative AI, a new concept is emerging: Agentic AI. While generative AI produces output in response to prompts, agentic AI systems promise to perform mundane tasks autonomously, potent…

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

2024-10-09 · Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung 외

We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering. To this end, we curate 75 ML engineering-related competitions from Kaggle, creating a diverse set of challengi…

TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents

2025-05-19 · Yifu Cai, Xinyu Li, Mononito Goswami, Michał Wiliński 외

We introduce TimeSeriesGym, a scalable benchmarking framework for evaluating Artificial Intelligence (AI) agents on time series machine learning engineering challenges. Existing benchmarks lack scalability, focus narrowl…

AI AgentBenchmarkingCode TranslationTime Series

Toward the Engineering of Virtuous Machines

2018-12-07 · Naveen Sundar Govindarajulu, Selmer Bringsjord, Rikhiya Ghosh

While various traditions under the 'virtue ethics' umbrella have been studied extensively and advocated by ethicists, it has not been clear that there exists a version of virtue ethics rigorous enough to be a target for …

EthicsFormal Logic