paper-with-me

홈 › Papers

Hierarchical Experimentalist Agents

2026-06-28 · Abhranil Chandra, Sankaran Vaidyanathan, Utsav Dhanuka, Varun Gandhi, Scott Niekum hf

Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-training data, retrieval, or search. This paradigm breaks down in novel domains and for sophisticated queries that cannot be answered from prior knowledge alone. Knowing the laws of physics, for instance, does not by itself enable LLMs to answer queries or complete long-horizon tasks in a complex physical system. To address this, we introduce Hierarchical Experimentalist Agents (HExA), an in-context self-improvement framework to learn from active experimentation. HExA iteratively designs and refines query-relevant experiments, learns a reusable library of composable skills from experience, and integrates experimental evidence to answer queries or take actions. HExA is training-free, compatible with any black-box model, and does not require external supervision, oracles, or offline data. To evaluate active experimentation, we introduce Interphyre, a tool-calling benchmark built on the PHYRE 2D procedural physics environment, where agents propose interventions and test hypotheses through simulation APIs. Experiments show that current LLM agents struggle in these settings, especially on the hardest levels of Interphyre. Claude Sonnet 4.6 achieves only 2% success, while HExA improves the same model to up to 77% success. HExA also improves open-weight models and outperforms agentic baselines such as ReAct and Reflexion. Moreover, using only skills learned from easier levels and transferred without active experimentation, HExA achieves 44% success, demonstrating the reusability and generalization of its learned skills. Overall, HExA shows that learning through active experimentation can help agents discover useful knowledge, acquire reusable skills, and make efficient progress on novel long-horizon tasks.

📄 PDF Abstract BibTeX arXiv:2606.29315

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Agents for self-driving laboratories applied to quantum computing

2024-12-10 · Shuxiang Cao, Zijian Zhang, Mohammed Alghadeer, Simone D Fasciati 외

Fully automated self-driving laboratories are promising to enable high-throughput and large-scale scientific discovery by reducing repetitive labour. However, effective automation requires deep integration of laboratory …

Language ModelingLanguage ModellingLarge Language Modelscientific discovery

Active Causal Experimentalist (ACE): Learning Intervention Strategies via Direct Preference Optimization

2026-02-02 · Patrick Cooper, Alvaro Velasquez arxiv

Discovering causal relationships requires controlled experiments, but experimentalists face a sequential decision problem: each intervention reveals information that should inform what to try next. Traditional approaches…

Domain Adaptation

Hierarchically-coupled hidden Markov models for learning kinetic rates from single-molecule data

2013-05-15 · Jan-Willem van de Meent, Jonathan E. Bronson, Frank Wood, Ruben L. Gonzalez Jr. 외

We address the problem of analyzing sets of noisy time-varying signals that all report on the same process but confound straightforward analyses due to complex inter-signal heterogeneities and measurement artifacts. In p…

Time SeriesTime Series Analysis

Qumus: Realization of An Embodied AI Quantum Material Experimentalist

2026-05-18 · Lihan Shi, Zhaoyi Joy Zheng, Xinzhe Juan, Yimin Wang 외 arxiv

While modern Large Language Models (LLMs) and agentic artificial intelligence (AI) have demonstrated transformative capabilities in digital domains, the realization of embodied AI capable of real-world scientific discove…

NoisyCUR: An algorithm for two-cost budgeted matrix completion

2021-04-16 · Dong Hu, Alex Gittens, Malik Magdon-Ismail

Matrix completion is a ubiquitous tool in machine learning and data analysis. Most work in this area has focused on the number of observations necessary to obtain an accurate low-rank approximation. In practice, however,…

Matrix CompletionVocal Bursts Valence Prediction