paper-with-me

홈 › Papers

Team of Thoughts: Efficient Test-time Scaling of Agentic Systems through Orchestrated Tool Calling

2026-02-18 · Jeffrey T. H. Wong, Zixi Zhang, Junyi Liu, Yiren Zhao arxiv

Existing Multi-Agent Systems (MAS) typically rely on homogeneous model configurations, failing to exploit the diverse expertise inherent in different post-trained architectures. We propose Team-of-Thoughts, a heterogeneous MAS framework that treats diverse models as specialized tools within an orchestrator-driven paradigm. Team-of-Thoughts introduces two novel components: (1) Orchestrator Calibration, which identifies models with superior coordination and synthesis capabilities, and (2) Agent Self-Assessment, a protocol where tool agents profile their own domain-specific strengths to guide selection. At inference, the orchestrator dynamically activates the most compatible agents based on these profiles to maximize capability coverage. Across five mathematical reasoning and code generation benchmarks, Team-of-Thoughts consistently outperforms individual models and existing MAS baselines. Notably, on AIME24 and LiveCodeBench, Team-of-Thoughts achieves 96.00% and 77.91% accuracy, respectively, significantly improving over homogeneous role-play baselines (80.00% and 65.93%).

📄 PDF Abstract BibTeX arXiv:2602.16485

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical ReasoningCode Generation

Similar Papers 제목 키워드 기반

OpenThoughts-Agent: Data Recipes for Agentic Models

2026-06-23 · Negin Raoof, Richard Zhuang, Marianna Nezhurina, Etash Guha 외 arxiv

Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Te…

Clover: A Neural-Symbolic Agentic Harness with Stochastic Tree-of-Thoughts for Verified RTL Repair

2026-04-19 · Zizhang Luo, Yansong Xu, Runlin Guo, Fan Cui 외 arxiv

RTL program repair remains a critical bottleneck in hardware design and verification. Traditional automatic program repair (APR) methods rely on predefined templates and synthesis, limiting their bug coverage. Large lang…

Program Repair

MetaScale: Test-Time Scaling with Evolving Meta-Thoughts

2025-03-17 · Qin Liu, Wenxuan Zhou, Nan Xu, James Y. Huang 외

One critical challenge for large language models (LLMs) for making complex reasoning is their reliance on matching reasoning patterns from training data, instead of proactively selecting the most appropriate cognitive st…

Diffuse Thinking: Exploring Diffusion Language Models as Efficient Thought Proposers for Reasoning

2025-10-31 · Chenyang Shao, Sijian Ren, Fengli Xu, Yong Li arxiv

In recent years, large language models (LLMs) have witnessed remarkable advancements, with the test-time scaling law consistently enhancing the reasoning capabilities. Through systematic evaluation and exploration of a d…

ARTIS: Agentic Risk-Aware Test-Time Scaling via Iterative Simulation

2026-02-02 · Xingshan Zeng, Lingzhi Wang, Weiwen Liu, Liangyou Li 외 arxiv

Current test-time scaling (TTS) techniques enhance large language model (LLM) performance by allocating additional computation at inference time, yet they remain insufficient for agentic settings, where actions directly …

Decision Making