paper-with-me

홈 › Papers

TinyClick: Single-Turn Agent for Empowering GUI Automation

2024-10-09 · Pawel Pawlowski, Krystian Zawistowski, Wojciech Lapacz, Adam Wiacek, Marcin Skorupa, Sebastien Postansque, Jakub Hoscilowicz

We present an UI agent for user interface (UI) interaction tasks, using Vision-Language Model Florence-2-Base. The agent's primary task is identifying the screen coordinates of the UI element corresponding to the user's command. It demonstrates very strong performance on Screenspot and OmniAct annotations, while maintaining a very small size of 0.27B parameters and minimal latency. Moreover, training needs small compute budget of 56 GPU-hours (worth about 40 USD). Relevant improvement comes from vision-specific multi-task training and MLLM-based data augmentation. We hope that decreased needs for expensive compute resources and manually annotated data will allow to facilitate more inclusive and sustainable research of UI agents.

📄 PDF Abstract BibTeX arXiv:2410.11871

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationGPULanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

AgentTutor: Empowering Personalized Learning with Multi-Turn Interactive Teaching in Intelligent Education Systems

2025-12-24 · Yuxin Liu, Zeqing Song, Jiong Lou, Chentao Wu 외 arxiv

The rapid advancement of large-scale language models (LLMs) has shown their potential to transform intelligent education systems (IESs) through automated teaching and learning support applications. However, current IESs …

Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?

2026-06-09 · Tengchao Lv, Dongdong Zhang, Jiayu Ding, Yilin Jia 외 arxiv

The deployment of Large Language Model (LLM) agents for computer automation is accelerating, yet their ability to navigate complex, professional-grade productivity software is largely untested. We argue that Office autom…

Code Generation

MedSAM-Agent: Empowering Interactive Medical Image Segmentation with Multi-turn Agentic Reinforcement Learning

2026-02-03 · Shengyuan Liu, Liuxin Bao, Qi Yang, Wanting Geng 외 arxiv

Medical image segmentation is evolving from task-specific models toward generalizable frameworks. Recent research leverages Multi-modal Large Language Models (MLLMs) as autonomous agents, employing reinforcement learning…

Medical Image SegmentationInteractive SegmentationReinforcement Learning

EvoFlow: Evolving Diverse Agentic Workflows On The Fly

2025-02-11 · Guibin Zhang, Kaijie Chen, Guancheng Wan, Heng Chang 외

The past two years have witnessed the evolution of large language model (LLM)-based multi-agent systems from labor-intensive manual design to partial automation (\textit{e.g.}, prompt engineering, communication topology)…

Large Language ModelPrompt EngineeringTAG

Multi-lingual Multi-turn Automated Red Teaming for LLMs

2025-04-04 · Abhishek Singhania, Christophe Dupuy, Shivam Mangale, Amani Namboori

Language Model Models (LLMs) have improved dramatically in the past few years, increasing their adoption and the scope of their capabilities over time. A significant amount of work is dedicated to ``model alignment'', i.…

Red Teaming