paper-with-me

홈 › Papers

Structuring GUI Elements through Vision Language Models: Towards Action Space Generation

2025-08-22 · Yi Xu, Yesheng Zhang, Jiajia Liu, Jingdong Chen arxiv

Multimodal large language models (MLLMs) have emerged as pivotal tools in enhancing human-computer interaction. In this paper we focus on the application of MLLMs in the field of graphical user interface (GUI) elements structuring, where they assist in processing user instructions based on screen contents. Despite the promise of MLLMs, their performance in precisely generating UI element coordinates, a critical aspect of GUI understanding, is hindered by the nature of next-token prediction training. This challenge arises from the semantic void surrounding numerical UI coordinates in language representation spaces, necessitating a substantial and diverse dataset to bolster visual module capabilities. To address these limitations, we introduce an IoU-Augmented Maximum Likelihood (IAML) training paradigm. Specifically, our approach involves a novel pipeline for IoU-based coordinate sampling to augment the training data, which considers the proximity to ground truth coordinates. This data augmentation strategy is then employed to fine-tune MLLMs under the IAML paradigm, which is designed to mitigate the exposure bias problem inherent in traditional maximum likelihood estimation. Through extensive experiments, we demonstrate the superior performance of our IAML training approach over traditional training paradigms.

📄 PDF Abstract BibTeX arXiv:2508.16271

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

Morphological Networks for Image De-raining

2019-01-08 · Ranjan Mondal, Pulak Purkait, Sanchayan Santra, Bhabatosh Chanda

Mathematical morphological methods have successfully been applied to filter out (emphasize or remove) different structures of an image. However, it is argued that these methods could be suitable for the task only if the …

SSIM

Morpho-logic from a Topos Perspective: Application to symbolic AI

2023-03-08 · Marc Aiguier, Isabelle Bloch, Salim Nibouche, Ramon Pino Perez

Modal logics have proved useful for many reasoning tasks in symbolic artificial intelligence (AI), such as belief revision, spatial reasoning, among others. On the other hand, mathematical morphology (MM) is a theory for…

Spatial Reasoning

Facilitating Self-Guided Mental Health Interventions Through Human-Language Model Interaction: A Case Study of Cognitive Restructuring

2023-10-24 · ASHISH SHARMA, Kevin Rushton, Inna Wanyin Lin, Theresa Nguyen 외

Self-guided mental health interventions, such as "do-it-yourself" tools to learn and practice coping strategies, show great promise to improve access to mental health care. However, these interventions are often cognitiv…

Language ModelingLanguage Modelling

Brief2Design: A Multi-phased, Compositional Approach to Prompt-based Graphic Design

2026-04-13 · Kotaro Kikuchi, Nami Ogawa arxiv

Professional designers work from client briefs that specify goals and constraints but often lack concrete design details. Translating these abstract requirements into visual designs poses a central challenge, yet existin…

Modular Framework for Visuomotor Language Grounding

2021-09-05 · Kolby Nottingham, Litian Liang, Daeyun Shin, Charless C. Fowlkes 외

Natural language instruction following tasks serve as a valuable test-bed for grounded language and robotics research. However, data collection for these tasks is expensive and end-to-end approaches suffer from data inef…

Instruction Following