paper-with-me

홈 › Papers

CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing

2026-05-19 · Haobo Hu, Xiangwu Guo, Zhiheng Chen, Difei Gao, Haotian Liu, Libiao Jin, Qi Mao arxiv

While GUI agents have made significant progress in web navigation and basic operating system tasks, their capabilities in professional creative workflows remain largely underexplored. To bridge this gap, we introduce Cutverse, a benchmark designed to systematically evaluate autonomous GUI agents in realistic media post-production environments. We curate expert demonstrations across 7 professional applications (e.g., Premiere Pro, Photoshop), covering 186 complex, long-horizon tasks grounded in authentic editing workflows, involving dense multimodal interfaces and tightly coupled interaction sequences. To support scalable evaluation, we develop a lightweight parser that transforms raw screen recordings and low-level interaction logs into structured, compositional GUI action trajectories with precise grounding. Extensive evaluations reveal that existing agents achieve only 36.0\% task success on realistic media editing tasks, underscoring the challenges posed by complex, long-horizon media post-production workflows in our benchmark.While current models demonstrate promising spatial grounding, multimodal alignment, and coordinated action execution, they remain limited in long-horizon reliability and domain-specific planning.

📄 PDF Abstract BibTeX arXiv:2605.19484

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

2026-07-29 · Lehan Wang, Boli Chen, Ruixue Ding, Pengjun Xie 외 arxiv

Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabi…

SoMe: A Realistic Benchmark for LLM-based Social Media Agents

2025-12-09 · Dizhan Xue, Jing Cui, Shengsheng Qian, Chuanrui Hu 외 arxiv

Intelligent agents powered by large language models (LLMs) have recently demonstrated impressive capabilities and gained increasing popularity on social media platforms. While LLM agents are reshaping the ecology of soci…

MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model

2026-03-19 · Youngwan Lee, Soojin Jang, Yoorhim Cho, Seunghwan Lee 외 arxiv

Spatial reasoning is foundational for Vision-Language Models (VLMs), particularly when deployed as Vision-Language-Action (VLA) agents in physical environments. However, existing benchmarks predominantly focus on element…

Reinforcement LearningSpatial ReasoningAnswer SelectionVisual Grounding

MMInA: Benchmarking Multihop Multimodal Internet Agents

2024-04-15 · Ziniu Zhang, Shulin Tian, Liangyu Chen, Ziwei Liu

Autonomous embodied agents live on an Internet of multimedia websites. Can they hop around multimodal websites to complete complex user tasks? Existing benchmarks fail to assess them in a realistic, evolving environment …

Benchmarking

Owner-Harm: A Missing Threat Model for AI Agent Safety

2026-04-20 · Dongcheng Zhang, Yiqing Jiang arxiv

Existing AI agent safety benchmarks focus on generic criminal harm (cybercrime, harassment, weapon synthesis), leaving a systematic blind spot for a distinct and commercially consequential threat category: agents harming…