paper-with-me

홈 › Papers

SCUBA: Salesforce Computer Use Benchmark

2025-09-30 · Yutong Dai, Krithika Ramakrishnan, Jing Gu, Matthew Fernandez, Yanqi Luo, Viraj Prabhu, Zhenyu Hu, Silvio Savarese, Caiming Xiong, Zeyuan Chen, Ran Xu arxiv

We introduce SCUBA, a benchmark designed to evaluate computer-use agents on customer relationship management (CRM) workflows within the Salesforce platform. SCUBA contains 300 task instances derived from real user interviews, spanning three primary personas, platform administrators, sales representatives, and service agents. The tasks test a range of enterprise-critical abilities, including Enterprise Software UI navigation, data manipulation, workflow automation, information retrieval, and troubleshooting. To ensure realism, SCUBA operates in Salesforce sandbox environments with support for parallel execution and fine-grained evaluation metrics to capture milestone progress. We benchmark a diverse set of agents under both zero-shot and demonstration-augmented settings. We observed huge performance gaps in different agent design paradigms and gaps between the open-source model and the closed-source model. In the zero-shot setting, open-source model powered computer-use agents that have strong performance on related benchmarks like OSWorld only have less than 5\% success rate on SCUBA, while methods built on closed-source models can still have up to 39% task success rate. In the demonstration-augmented settings, task success rates can be improved to 50\% while simultaneously reducing time and costs by 13% and 16%, respectively. These findings highlight both the challenges of enterprise tasks automation and the promise of agentic solutions. By offering a realistic benchmark with interpretable evaluation, SCUBA aims to accelerate progress in building reliable computer-use agents for complex business software ecosystems.

📄 PDF Abstract BibTeX arXiv:2509.26506

Code (0)

등록된 구현이 없습니다.

Tasks

Information Retrieval

Similar Papers 제목 키워드 기반

Robotic Classification of Divers' Swimming States using Visual Pose Keypoints as IMUs

2025-10-15 · Demetrious T. Kutzke, Ying-Kun Wu, Elizabeth Terveen, Junaed Sattar arxiv

Traditional human activity recognition uses either direct image analysis or data from wearable inertial measurement units (IMUs), but can be ineffective in challenging underwater environments. We introduce a novel hybrid…

Human Activity Recognition

Synchronized SCUBA: D2D Communication for Out-of-Sync Devices

2021-04-01 · Vishnu Rajendran Chandrika, Gautham Prasad, Lutz Lampe, Gus Vos

Device-to-device (D2D) communication is an essential component enabling connectivity for the Internet-of-Things (IoT). SCUBA, which stands for Sidelink Communication on Unlicensed Bands, is a novel medium access control …

SCUBA: An In-Device Multiplexed Protocol for Sidelink Communication on Unlicensed Bands

2020-12-07 · Vishnu Rajendran, Gautham Prasad, Lutz Lampe, Gus Vos

Device-to-device communication (D2D) is a key enabler for connecting devices together to form the Internet of Things (IoT). A growing issue with IoT networks is the increasing number of IoT devices congesting the spectra…

Latest News in Computational Argumentation: Surfing on the Deep Learning Wave, Scuba Diving in the Abyss of Fundamental Questions

2017-09-01 · WS 2017 9 · Iryna Gurevych

Mining arguments from natural language texts, parsing argumentative structures, and assessing argument quality are among the recent challeng-es tackled in computational argumentation. While advanced deep learning models …

Common Sense Reasoning

BrainSCUBA: Fine-Grained Natural Language Captions of Visual Cortex Selectivity

2023-10-06 · Andrew F. Luo, Margaret M. Henderson, Michael J. Tarr, Leila Wehbe

Understanding the functional organization of higher visual cortex is a central focus in neuroscience. Past studies have primarily mapped the visual and semantic selectivity of neural populations using hand-selected stimu…

Image GenerationLanguage ModelingLanguage ModellingLarge Language Model+1