paper-with-me

홈 › Papers

VESTA: Visual Exploration with Statistical Tool Agents

2026-05-29 · William Rudman, Abhishek Divekar, Kanishk Jain, Sebastian Joseph, Stella S. R. Offner, Matthew Lease, Kyle Mahowald, Greg Durrett, Junyi Jessy Li arxiv

Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated. Recent agent-based systems leverage language and vision-language models (VLMs) to iteratively propose and refine statistical models, but these systems struggle on more challenging modeling tasks. To address these limitations, we introduce VESTA: Visual Exploration with Statistical Tool Agents, a framework that equips VLMs with a dynamically growing exploration toolkit to guide model refinement through data transformations, hypothesis-driven visualizations, and robust statistical tests. Unlike prior systems that rely on iterative critique alone, VESTA actively explores data before and during refinement by selecting or creating diagnostic tools, which accumulate in the model's context and can be reused later. We evaluate VESTA against established baselines in three toolkit configurations: no tools, static expert-written tools, and dynamic model-written tools. To support this evaluation, we introduce DAWN (Dataset for Automated Workflows and Numerical Modeling), a benchmark targeting distribution fitting and time series modeling with varying difficulty tiers, and culminating in real-world astronomy tasks including modeling initial mass functions and gravitational-wave chirp signals. We find that VESTA's dynamic tool creation outperforms prior agentic pipelines, with the largest gains on complex and domain-specific tasks. We further show that dynamically generated tools are substantially more sophisticated than those produced by existing visual tool-creation systems, covering more diagnostic categories per function and strongly preferring visual outputs that the VLM critic can reason over directly.

📄 PDF Abstract BibTeX arXiv:2606.00384

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VESTA: A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

2026-06-07 · Lu Jia, Haibo Tong, Feifei Zhao, Jindong Li 외 arxiv

Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks. As their capabilitie…

From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents

2026-08-31 · Can Zhang, Baofeng Zhang, Xiaotian Han, Junyuan Shang 외 arxiv

Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurrence coverage, or must discriminate competing hypotheses---which can …

DriveStack-VLA: Render-Teacher Alignment for BEV-Based DeepStack Vision-Language-Action Model

2026-06-23 · Jingke Wang, Zhenru Zhao, Shuangming Lei, Hao Su 외 arxiv

Vision-Language-Action driving models convert a pretrained Vision-Language Model into a driving policy, allowing them to use world knowledge and follow language guidances. However, existing VLA driving models still lack …

Motion Planning

Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents

2026-05-25 · Dong-Hee Kim, Reuben Tan, Donghyun Kim arxiv

Visual agents employ external visual tools within visual chains of thought to incorporate fine-grained evidence. While prior work has mainly studied these tools in visual search tasks, their role in more complex visual r…

Visual Question AnsweringSpatial ReasoningVisual Reasoning

InvestAlign: Overcoming Data Scarcity in Aligning Large Language Models with Investor Decision-Making Processes under Herd Behavior

2025-07-09 · Huisheng Wang, Zhuoshi Pan, Hangjing Zhang, Mingxiao Liu 외

Aligning Large Language Models (LLMs) with investor decision-making processes under herd behavior is a critical challenge in behavioral finance, which grapples with a fundamental limitation: the scarcity of real-user dat…

Decision Making