paper-with-me

홈 › Papers

MACEval: A Multi-Agent Continual Evaluation Network for Large Models

2025-11-12 · Zijian Chen, Yuze Sun, Yuan Tian, Wenjun Zhang, Guangtao Zhai arxiv

Hundreds of benchmarks dedicated to evaluating large models have been presented over the past few years. However, most of them remain closed-ended and are prone to overfitting due to the potential data contamination. Moreover, the increasing scale and scope of current benchmarks with transient metrics, as well as the heavily human-dependent curation procedure, pose significant challenges for timely maintenance and adaptation. In this paper, we introduce MACEval, a Multi-Agent Continual Evaluation network for dynamic evaluation of large models, and define new metrics to quantify performance longitudinally. MACEval employs an interactive and autonomous evaluation mode, utilizing role assignment, in-process data generation, and evaluation routing through a cascaded agent network. Extensive experiments on 23 large models demonstrate the effectiveness of MACEval, which also lightens the evaluation process and reduces a considerable amount of overhead. We hope that MACEval can broaden future directions of large model evaluation. Project page: https://github.com/zijianchen98/MACEval.

📄 PDF Abstract BibTeX arXiv:2511.09139

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Continual Reinforcement Learning with TELLA

2022-08-08 · Neil Fendley, Cash Costello, Eric Nguyen, Gino Perrotta 외

Training reinforcement learning agents that continually learn across multiple environments is a challenging problem. This is made more difficult by a lack of reproducible experiments and standard metrics for comparing di…

Continual LearningLifelong learningreinforcement-learningReinforcement Learning+1

SWE-Bench-CL: Continual Learning for Coding Agents

2025-06-13 · Thomas Joshi, Shayan Chowdhury, Fatih Uysal

Large Language Models (LLMs) have achieved impressive results on static code-generation benchmarks, but real-world software development unfolds as a continuous stream of evolving issues, fixes, and feature requests. We i…

Code GenerationContinual Learning

Vision-Language Navigation with Continual Learning

2024-09-04 · Zhiyuan Li, Yanfeng Lv, Ziqin Tu, Di Shang 외

Vision-language navigation (VLN) is a critical domain within embedded intelligence, requiring agents to navigate 3D environments based on natural language instructions. Traditional VLN research has focused on improving e…

Continual LearningNavigateVision-Language Navigation

AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents

2026-06-01 · Yiheng Shu, Bernal Jiménez Gutiérrez, Saisri Padmaja Jonnalagedda, Yuguang Yao 외 arxiv

Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes. Continual learning expects an agent to accumulate reusable …

Continual Learning

Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning

2026-01-07 · Zheng Wu, Xingyu Lou, Xinbei Ma, Yansi Li 외 arxiv

Large Language Model (LLM)-based agents significantly extend the utility of LLMs by interacting with dynamic environments. However, enabling agents to continually learn new tasks without catastrophic forgetting remains a…

Continual Learning