paper-with-me

홈 › Papers

FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks

2026-04-07 · Michael Krumdick, Varshini Reddy, Shivani Chaudhary, William Day, Maarij Ahmed, Hayan Haqqi, Muhammad Ahsen Fahim, Hanzallah Amjad, Ahmad Orakzai, Aqsa Gul, Chris Tanner arxiv

As concerns surrounding AI-driven labor displacement intensify in knowledge-intensive sectors, existing benchmarks fail to measure performance on tasks that define practical professional expertise. Finance, in particular, has been identified as a domain with high AI exposure risk, yet lacks robust benchmarks to track real-world developments. This gap is compounded by the absence of clear accountability mechanisms in current Large Language Model (LLM) deployments. To address this, we introduce FrontierFinance, a long-horizon benchmark of 25 complex financial modeling tasks across five core finance models, requiring an average of over 18 hours of skilled human labor per task to complete. Developed with financial professionals, the benchmark reflects industry-standard financial modeling workflows and is paired with detailed rubrics for structured evaluation. We engage human experts to define the tasks, create rubrics, grade LLMs, and perform the tasks themselves as human baselines. We demonstrate that our human experts both receive higher scores on average, and are more likely to provide client-ready outputs than current state-of-the-art systems.

📄 PDF Abstract BibTeX arXiv:2604.05912

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

2026-06-08 · Wanli Li, Bowen Zhou, Yunyao Yu, Zhou Xu 외 arxiv

Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools. Existing benchmarks, however, often evaluate these inte…

LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios

2026-02-03 · Tianyu Chen, Chujia Hu, Ge Gao, Dongrui Liu 외 arxiv

Computer-use agents (CUAs) that interact with real computer systems can perform automated tasks but face critical safety risks. Ambiguous instructions may trigger harmful actions, and adversarial users can manipulate too…

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

2026-06-28 · Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong, Weiming Wu 외 arxiv

Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents. We introduce OSWorld 2.0…

CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design

2026-09-14 · Zihan Dong, Yuanzhe Liu, Zhiyuan Ma, Qishi Zhan 외 hf

Computer-use agents are increasingly evaluated in realistic desktop environments, but existing benchmarks provide limited coverage of professional engineering workflows whose outputs are persistent, structured artifacts.…

CarePilot: A Multi-Agent Framework for Long-Horizon Computer Task Automation in Healthcare

2026-03-25 · Akash Ghosh, Tajamul Ashraf, Rishu Kumar Singh, Numan Saeed 외 arxiv

Multimodal agentic pipelines are transforming human-computer interaction by enabling efficient and accessible automation of complex, real-world tasks. However, recent efforts have focused on short-horizon or general-purp…