paper-with-me

Papers

MAPS: A Multilingual Benchmark for Global Agent Performance and Security

2025-05-21 · Omer Hofman, Oren Rachmil, Shamik Bose, Vikas Pahuja, Jonathan Brokman, Toshiya Shimizu, Trisha Starostina, Kelly Marchisio, Seraphina Goldfarb-Tarrant, Roman Vainshtein

Agentic AI systems, which build on Large Language Models (LLMs) and interact with tools and memory, have rapidly advanced in capability and scope. Yet, since LLMs have been shown to struggle in multilingual settings, typically resulting in lower performance and reduced safety, agentic systems risk inheriting these limitations. This raises concerns about the global accessibility of such systems, as users interacting in languages other than English may encounter unreliable or security-critical agent behavior. Despite growing interest in evaluating agentic AI, existing benchmarks focus exclusively on English, leaving multilingual settings unexplored. To address this gap, we propose MAPS, a multilingual benchmark suite designed to evaluate agentic AI systems across diverse languages and tasks. MAPS builds on four widely used agentic benchmarks - GAIA (real-world tasks), SWE-bench (code generation), MATH (mathematical reasoning), and the Agent Security Benchmark (security). We translate each dataset into ten diverse languages, resulting in 805 unique tasks and 8,855 total language-specific instances. Our benchmark suite enables a systematic analysis of how multilingual contexts affect agent performance and robustness. Empirically, we observe consistent degradation in both performance and security when transitioning from English to other languages, with severity varying by task and correlating with the amount of translated input. Building on these findings, we provide actionable recommendations to guide agentic AI systems development and assessment under multilingual settings. This work establishes a standardized evaluation framework, encouraging future research towards equitable, reliable, and globally accessible agentic AI. MAPS benchmark suite is publicly available at https://huggingface.co/datasets/Fujitsu-FRE/MAPS

📄 PDF Abstract BibTeX arXiv:2505.15935

Code (0)

등록된 구현이 없습니다.

Tasks

Code GenerationMathMathematical Reasoning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MASim: Multilingual Agent-Based Simulation for Social Science

2025-12-08 · Xuan Zhang, Wenxuan Zhang, Anxu Wang, See-Kiong Ng 외 arxiv

Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly monolingual and fail to model cross-lingual interaction, an essential property of…

GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation

2026-04-27 · Yunsu Kim, Kaden Uhlig, Joern Wuebker arxiv

Agent benchmarks remain largely English-centric, while their multilingual versions are often built with machine translation (MT) and limited post-editing. We argue that, for agentic tasks, this minimal workflow can easil…

Machine Translation

X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System

2025-05-21 · Peng Wang, Ruihan Tao, Qiguang Chen, Mengkang Hu 외

Recently, large language model (LLM)-based agents have achieved significant success in interactive environments, attracting significant academic and industrial attention. Despite these advancements, current research pred…

Language ModelingLanguage ModellingLarge Language Model

PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts

2026-05-13 · Yifei Zhu arxiv

Large Reasoning Models (LRMs) embedded in agentic frameworks have transformed information retrieval from static, long context question answering into open-ended exploration. Yet real world use requires models to discover…

Information RetrievalQuestion Answering

DeepMNavigate: Deep Reinforced Multi-Robot Navigation Unifying Local & Global Collision Avoidance

2019-10-04 · Qingyang Tan, Tingxiang Fan, Jia Pan, Dinesh Manocha

We present a novel algorithm (DeepMNavigate) for global multi-agent navigation in dense scenarios using deep reinforcement learning (DRL). Our approach uses local and global information for each robot from motion informa…

Collision AvoidanceDeep Reinforcement LearningPositionreinforcement-learning+3