paper-with-me

Papers

GUI Odyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

2024-06-12 · Quanfeng Lu, Wenqi Shao, Zitao Liu, Fanqing Meng, Boxuan Li, Botong Chen, Siyuan Huang, Kaipeng Zhang, Yu Qiao, Ping Luo

Smartphone users often navigate across multiple applications (apps) to complete tasks such as sharing content between social media platforms. Autonomous Graphical User Interface (GUI) navigation agents can enhance user experience in communication, entertainment, and productivity by streamlining workflows and reducing manual intervention. However, prior GUI agents often trained with datasets comprising simple tasks that can be completed within a single app, leading to poor performance in cross-app navigation. To address this problem, we introduce GUI Odyssey, a comprehensive dataset for training and evaluating cross-app navigation agents. GUI Odyssey consists of 7,735 episodes from 6 mobile devices, spanning 6 types of cross-app tasks, 201 apps, and 1.4K app combos. Leveraging GUI Odyssey, we developed OdysseyAgent, a multimodal cross-app navigation agent by fine-tuning the Qwen-VL model with a history resampling module. Extensive experiments demonstrate OdysseyAgent's superior accuracy compared to existing models. For instance, OdysseyAgent surpasses fine-tuned Qwen-VL and zero-shot GPT-4V by 1.44\% and 55.49\% in-domain accuracy, and 2.29\% and 48.14\% out-of-domain accuracy on average. The dataset and code will be released in \url{https://github.com/OpenGVLab/GUI-Odyssey}.

📄 PDF Abstract BibTeX arXiv:2406.08451

Code (1)

opengvlab/gui-odyssey 공식 구현 pytorch

Tasks

Navigate

Similar Papers 제목 키워드 기반

Odyssey: An Automotive Lidar-Inertial Odometry Dataset with GNSS-denied situations

2025-12-16 · Aaron Kurda, Simon Steuernagel, Lukas Jung, Marcus Baum arxiv

The development and evaluation of Lidar-Inertial Odometry (LIO) and Simultaneous Localization and Mapping (SLAM) systems requires a precise ground truth. The Global Navigation Satellite System (GNSS) is often used as a f…

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

2026-05-21 · Haichen He, Jiayi Zhou, Sifeng Shang, Yihan Hu 외 arxiv

Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal spans within extreme video durations. Mastering this intense cognitiv…

Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks

2026-04-27 · Lawrence Keunho Jang, Jing Yu Koh, Daniel Fried, Ruslan Salakhutdinov arxiv

Existing web agent benchmarks have largely converged on short, single-site tasks that frontier models are approaching saturation on. However, real world web use consists of long-horizon, multi-site workflows. Common web …

OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

2025-08-12 · Weixuan Wang, Dongge Han, Daniel Madrigal Diaz, Jin Xu 외 arxiv

Autonomous agents powered by large language models (LLMs) are increasingly deployed in real-world applications requiring complex, long-horizon workflows. However, existing benchmarks predominantly focus on atomic tasks t…

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents

2025-05-19 · CVPR 2025 1 · Yunseok Jang, Yeda Song, Sungryull Sohn, Lajanugen Logeswaran 외

Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have sparked significant interest in developing GUI visual agents. We introduce MONDAY (Mobile OS Navigation Task Dataset for Agents f…

Dataset GenerationOptical Character Recognition (OCR)