paper-with-me

Papers

X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System

2025-05-21 · Peng Wang, Ruihan Tao, Qiguang Chen, Mengkang Hu, Libo Qin

Recently, large language model (LLM)-based agents have achieved significant success in interactive environments, attracting significant academic and industrial attention. Despite these advancements, current research predominantly focuses on English scenarios. In reality, there are over 7,000 languages worldwide, all of which demand access to comparable agentic services. Nevertheless, the development of language agents remains inadequate for meeting the diverse requirements of multilingual agentic applications. To fill this gap, we introduce X-WebAgentBench, a novel multilingual agent benchmark in an interactive web environment, which evaluates the planning and interaction performance of language agents across multiple languages, thereby contributing to the advancement of global agent intelligence. Additionally, we assess the performance of various LLMs and cross-lingual alignment methods, examining their effectiveness in enhancing agents. Our findings reveal that even advanced models like GPT-4o, when combined with cross-lingual techniques, fail to achieve satisfactory results. We hope that X-WebAgentBench can serve as a valuable benchmark for multilingual agent scenario in real-world applications.

📄 PDF Abstract BibTeX arXiv:2505.15372

Code (1)

WPENGxs/X-WebAgentBench 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

2024-10-09 · Ido Levy, Ben wiesel, Sami Marreed, Alon Oved 외

Autonomous web agents solve complex browsing tasks, yet existing benchmarks measure only whether an agent finishes a task, ignoring whether it does so safely or in a way enterprises can trust. To integrate these agents i…

Autonomous Web Navigation

macOSWorld: A Multilingual Interactive Benchmark for GUI Agents

2025-06-04 · Pei Yang, Hai Ci, Mike Zheng Shou

Graphical User Interface (GUI) agents show promising capabilities for automating computer-use tasks and facilitating accessibility, but existing interactive benchmarks are mostly English-only, covering web-use or Windows…

BenchmarkingDomain Adaptation

MAPS: A Multilingual Benchmark for Global Agent Performance and Security

2025-05-21 · Omer Hofman, Oren Rachmil, Shamik Bose, Vikas Pahuja 외

Agentic AI systems, which build on Large Language Models (LLMs) and interact with tools and memory, have rapidly advanced in capability and scope. Yet, since LLMs have been shown to struggle in multilingual settings, typ…

Code GenerationMathMathematical Reasoning

Evaluating Large Language Model with Knowledge Oriented Language Specific Simple Question Answering

2025-05-22 · Bowen Jiang, Runchuan Zhu, Jiang Wu, Zinco Jiang 외

We introduce KoLasSimpleQA, the first benchmark evaluating the multilingual factual ability of Large Language Models (LLMs). Inspired by existing research, we created the question set with features such as single knowled…

Global FactsLanguage ModelingLanguage ModellingLarge Language Model+2

Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models

2025-07-02 · Samridhi Raj Sinha, Rajvee Sheth, Abhishek Upperwal, Mayank Singh arxiv

The rapid evolution of Large Language Models' has underscored the need for evaluation frameworks that are globally applicable, flexible, and modular, and that support a wide range of tasks, model types, and linguistic se…